DreamDB

Spec 0016 — Streaming Updates and Real-Time Freshness

Status: Draft (Phase 4 design). Depends on: spec/0001, spec/0004, spec/0006, spec/0008, spec/0010, spec/0013. Motivation: DreamDB's immutable + paged data plane is exquisite for archival and large-batch ingest, but the steady-state "ingest now, query in <1 second" workload exposes two structural gaps: (1) every write must produce a new Manifest, which costs a per-publish HTTP round-trip (typically 30–100 ms on commodity backends) — many small writes per second are wasteful; (2) spec/0013 explicitly defers FreshDiskANN (incremental graph updates), so continuous ingest past ~100M forces a full graph rebuild — impractical. Production retrieval systems all solve this with a hot delta tier in front of the cold base tier, plus a streaming graph-update algorithm, plus a drift-monitoring signal so operators know when to re-train. spec/0016 adds those three primitives.


1. Purpose

DreamDB's current "write a Manifest per batch, re-resolve from a ref" cycle is correct semantically but slow at high write rates and high query freshness. The fix is a per-Track HotShard Object that holds recent appends in a compact form, plus protocol-level signals for index health (when to trigger re-training) and an incremental-update algorithm for graph indexes.

By the end of this document the following are concrete:

  • The HotShard Object v1: a raw-f32 embedding buffer for one Track; v2 adds the typed scalar/text overlay of §2.7. Other field kinds remain refused.
  • The bounded-staleness contract: how readers control freshness vs latency tradeoff.
  • The dreamdb.fresh-vamana-cosine algorithm: streaming Vamana with append + background consolidation (FreshDiskANN, Singh et al. 2021).
  • The index-health signal: per-modality registry field declaring estimated recall, drift metric, last-train timestamp. Operators / SDKs use it to schedule re-training.
  • The Ada-IVF style background maintenance: when drift exceeds threshold, the SDK / operator triggers re-training; the new SpatialIndex/VectorCompressor Object lands via a Layer Manifest.

What stays defined elsewhere:

  • Per-modality index byte format — spec/0007.
  • VectorCompressor codebook publication — spec/0010.
  • Graph Index / Page format — spec/0013.

What this document does NOT define:

  • Strongly-consistent live reads. HotShard provides bounded-staleness, NOT linearizable freshness. Linearizable reads require cross-actor consensus, which spec/0008 explicitly does not provide.
  • Manifest publishing rate limits. That's a connector-level concern; spec/0006 already covers it.
  • Hot-tier durability guarantees beyond "fsync before commit." Backend-specific.

2. The HotShard Object

A HotShard is a per-Track buffer of recent Items, content-addressed like everything else.

2.1 Address path

<timeline>/<modality>/hot-shard/<hotshard-hash>

(New per-Timeline slot, parallel to track/.)

2.2 CBOR encoding (v1)

The v1 map has exactly the following semantic fields. Each items member is a two-element array; it is not a map. payload is exactly dim * 4 bytes of little-endian f32 values for the embedding field whose Track the registry entry binds. flush_threshold and ttl_seconds are producer policy values, not derived statistics.

{
  "version":         1,
  "parent_track":    <multihash>,                  ;; the immutable Track this overlay extends
  "published_at":    <u64>,                        ;; Unix ns at publish time; required for §3.3 staleness check
  "flush_threshold": <u32>,
  "ttl_seconds":     <u32>,
  "items": [                                       ;; ordered by time_anchor
    [<u64 anchor>, <bstr raw-f32le>],
    …
  ]
}

The Item count and earliest/latest anchors are derived from items; they are not duplicated on the wire. v1 has no spatial_keys, payload-kind tag, nested statistics, sub-Object reference, text value, or scalar value. A reader MUST reject an unsupported version rather than interpret it as v1. Adding another payload kind requires a new version whose encoding says how that kind is identified and read; opaque bytes alone are not a type system.

One HotShard overlays exactly one field-qualified embedding Track. The registry is keyed by that Track's allocated modality, so two fields MUST NOT share a modality key. A writer encountering such an ambiguous legacy binding MUST refuse instead of choosing one field or letting the buffers overwrite each other.

2.3 Registry reference

The Manifest registry's per-modality entry gains an optional hot_shard field:

"registry": {
  "embedding.f32.dim=768.bucketed.spatial_bits=18": {
    "kind":          "continuous",
    "object_kind":   "spatial-bucket",
    "algorithm":     "dreamdb.imi-cosine",
    "spatial_index": [<hash>],
    "track":         <multihash-Track-Object>,
    "hot_shard":     <multihash-HotShard | null>,   ;; NEW
  }
}

Absent hot_shard ⇒ Track behaves identically to v0 (cold-only). Present ⇒ readers MUST consult both Track and HotShard during query resolution.

2.4 Append path (writer)

A writer appending small/fast batches considers every embedding field present in each Sample. Each field has an independent v1 buffer and independent flush decision, but all updated HotShard references and cold Track replacements from one call become visible in one child Manifest:

1. Build the new Item(s).
2. Resolve each present embedding field to its unique allocated modality and
   fetch that field's current HotShard.
3. Append each field value to only that field's in-memory buffer.
4. For each field, if (`item_count >= flush_threshold`) OR
   (`age >= ttl_seconds`), independently:
     a. Build a replacement Track incorporating that field's buffered items.
     b. Clear that field's `hot_shard` reference in the same child Manifest.
   Else:
     a. Encode and PUT that field's new HotShard Object.
     b. Point that field's registry entry at it (Track Object unchanged).
5. Publish all changes from the call atomically. An unrelated field's live
   buffer remains referenced byte-for-byte when that field did not change.

Steps 1–3 + 4b are the hot path — no Track rewrite and no spatial index reorganization, just one small HotShard PUT per affected field and one ref CAS. This reduces the Objects rebuilt, not a guaranteed latency ratio. The reference writer evaluates age during append; ttl_seconds does not schedule a background timer. If writes stop, a published HotShard can remain referenced until an explicit flush. Readers and collectors must continue honoring it, including in retained ancestor Manifests; TTL is not expiration or permission to drop data.

2.5 Read path

A reader resolving a query for a field whose Track has a HotShard:

1. Resolve Manifest → Track Object + HotShard Object hashes.
2. Fetch both; their fetch order is not part of the format contract.
3. Track Object → identify cold candidate buckets / fragments.
4. HotShard v1 → validate raw-f32 width and produce vector candidates for this
   field only.
5. Merge cold + hot candidates; apply standard ranking (per-modality semantics).
6. Return merged top-K.

HotShard items live in CBOR, not in the modality's Spatial Bucket format — they bypass the spatial index. For vector ANN queries, the reader scores the field's HotShard vectors and merges them with the cold candidates under the same filtering, pool-selection and rerank policy. v1 does not define text or scalar query behavior.

The item count is bounded by the HotShard's flush_threshold parameter. Default: 10K items. v1 does not carry or enforce a byte threshold, so producers must size the item threshold with embedding dimension and memory cost in mind.

2.6 Forward-compat / backwards-compat

hot_shard is a critical extension under 0002 §3.1.0. It is not covered by the unknown-key hatch of 0002 §3.1.3, and the two paragraphs this section previously carried — describing a v0 reader as merely losing freshness — understated its effect on two independent axes.

  • Visibility (condition 2). An implementation that ignores hot_shard does not merely lose "freshness" in the sense of data not yet written. The buffered Items have been published: the HotShard Object was PUT and the Manifest naming it was committed by a ref CAS. Ignoring the field makes already-published data invisible to a query, and a reader cannot tell that state apart from the data never having been written. "Correct on what it does see" is not a useful guarantee when what it does not see was published.
  • Closure (condition 3). A HotShard that has not yet been folded into a Track is named only by this registry key. A collector that ignores the key does not mark the Object, and the sweep deletes it together with every buffered Item it holds.

Two operational facts bound any recovery plan, and neither is implied by the TTL:

  • The ttl_seconds forced flush of §3.1 is evaluated on the next append, not by a background timer. Halting writes does not drain an outstanding HotShard; only a further append, or an explicit flush, does.
  • A HotShard is not reachable solely from the newest Ref. Ancestor Manifests, other Refs and any retained root may continue to name an earlier HotShard Object, and a collector must mark it from each of those roots.

A v0.X reader (post-0016) seeing a Manifest without hot_shard reads exactly the v0 path — byte-identical hashes. That direction is unaffected by this section.

hot_shard is a named historical exception under 0002 §3.1.0.1: the registry key remains a valid format that MUST NOT be reinterpreted or require rewriting, an implementation declaring freshness support MUST implement the merge and marking semantics above, and an implementation without that capability MUST refuse explicitly rather than query or sweep against a Manifest that uses it. The per-role obligations, and the rollback rule assessed over every retained root a collector will process, are in 0002 §3.1.0.3.

2.7 Typed scalar/text HotShard v2

v1 embedding bytes and interpretation are unchanged. v2 supports the six existing Scalar field types, including UTF-8 String and Categorical text; it does not implement legacy FieldKind::Text, BM25 ingestion, media or arrays. Those kinds MUST be refused before publishing any part of a Sample.

The v2 canonical CBOR map is CLOSED, with exactly these seven keys:

{
  "version": 2,
  "parent_track": <bstr multihash of the exact cold Track>,
  "field_kind": "int" | "float" | "bool" | "string" | "categorical" | "timestamp",
  "published_at": <u64 Unix ns at buffer creation>,
  "flush_threshold": <positive u32>,
  "ttl_seconds": <u32>,
  "items": [[<u64 anchor>, <bstr payload>], ...]
}

Anchors MUST be strictly increasing and unique, and smaller than u64::MAX (the overlay uses a half-open extent). Int and Timestamp payloads are exactly eight little-endian signed i64 bytes. Float is exactly eight little-endian IEEE-754 binary64 bytes, finite only, not a CBOR float. Bool is exactly one byte, 0 or 1. String and Categorical are valid UTF-8 bytes, including empty text. The kind MUST match the bound Schema; a wrong kind, width, encoding, duplicate/unknown key, noncanonical order, or unknown version is refused. Payloads are inline primitive values and contain no downstream references.

Mandatory interpretation boundary. A scalar Track carrying hot data MUST use this CLOSED object_index map, not an ignorable optional key alone:

{"form": "hot-overlay-v1", "base_track": <bstr multihash>,
 "hot_shard": <bstr multihash>}

This form is valid only for ScalarBucket Tracks. The base MUST be a retained binding of the same timeline and modality in the Manifest's active/lineage closure, and decode as a cold scalar inline index or Scalar B-tree, never another overlay. Its hash MUST equal parent_track. The field-qualified registry hot_shard MUST equal the active overlay's hash (retained predecessor overlays keep their own historical hashes). The wrapper's coverage MUST include the base and every hot anchor. These are content-addressed references, not advisory metadata. Unsupported consumers MUST refuse this Track form before interpreting, rewriting, marking or sweeping it. No claim is made that every historically released implementation already obeys that requirement.

Append-only domain. On a nonempty cold base, every hot anchor MUST be at least the cold coverage's exclusive upper bound. An empty base imposes no lower bound. This is deliberately not historical backfill or scalar upsert. Within a hot batch/buffer, identical anchor and value bytes deduplicate; different values for one anchor are refused, never first/last-wins.

Publication and reads. Rust append_hot and its identity-aware adapters, Python append_hot/append_many(hot=True), and WASM Writer appendHot validate the whole Sample before publication. A single Manifest/Ref CAS publishes all changed field buffers and any threshold-triggered cold replacements. Admission or CAS failure must not mutate the caller handle to unpublished Track state. Concurrent writers receive the existing Ref conflict; no implicit retry or lost-update merge is introduced. Commit staged cold writes before hot append.

Reopening the Ref reads the exact wrapper, verified cold base and verified v2 buffer. Time scans, scalar predicates, facets and materialization combine both sources; tombstones suppress anchors in both. String/Categorical use existing scalar comparisons, not full-text search. iter_stream does not support typed hot scalar materialization and MUST refuse rather than silently return only cold values; use iter. Existing handles remain snapshot-bound until reopened.

Flush and maintenance. Producer threshold/TTL is evaluated on append; there is no background timer. Explicit flush_hot/flushHot materializes each value into the existing scalar cold writer and removes the registry overlay in the same publication. A second flush is a no-op. No intermediate cold-only Manifest is visible. Cold append, compaction and union merge with a pending typed scalar overlay MUST refuse and request explicit flush first; a fast forward may preserve an exact Manifest without rewriting its contents. This does not promise online union of independent hot scalar buffers.

Collectors follow and verify both wrapper references and the cold base's normal closure, including from retained ancestor roots. The lineage edge retains the base even after it is no longer query-active. Tombstones do not make referenced bytes collectible. Opaque byte copy remains permitted under 0002 §3.1.0; semantic copy must understand this form or refuse.

The v2 buffer is one Object, read and rewritten in full. Flush currently uses the ordinary in-memory scalar writer. This is not bounded-memory streaming, a per-Item-object format, a timer service, or a full-text indexing update.

3. The bounded-staleness contract

HotShard freshness has a producer-controlled upper bound and a consumer-controlled tolerance:

3.1 Producer side

The producer commits to a ttl_seconds (HotShard field). At the next hot append, or when an explicit flush operation runs, an older HotShard is forced into its cold Track even if its item count is below threshold. v1 does not define a background timer. Default: 30 seconds.

Smaller TTL ⇒ stronger freshness, more publish overhead. Larger TTL ⇒ weaker freshness, less overhead.

3.2 Consumer side

The consumer specifies a max_staleness_seconds in the HybridQuery (per spec/0015) or in a Query verb option:

  • If (now - hotshard.published_at) ≤ max_staleness_seconds: use the cached HotShard.
  • Else: re-resolve the ref, refresh the HotShard, then query.

This puts the freshness/cost tradeoff in the consumer's hands. A "must be current" query pays the round-trip; a "best effort" query takes the cached HotShard.

3.3 Latest-publish-at metadata

To support the consumer-side check, each HotShard Object carries a published_at: u64 (Unix ns) — set by the producer at publish time. This timestamp is non-authoritative for correctness (DreamDB content is time-anchored; freshness is a separate axis) but is the load-bearing signal for cache TTL.

The producer's published_at MUST be monotonically non-decreasing across HotShard publishes for the same Track — otherwise consumers' freshness checks misorder. Per spec/0008 monotonic-ts discipline.

4. Streaming Vamana — dreamdb.fresh-vamana-cosine

The spec/0013 §5 placeholder is filled. FreshDiskANN (Singh et al. 2021) defines an incremental Vamana that supports both append and delete without full rebuild.

4.1 Inheritance from spec/0013

Same GraphPage record and search algorithm. The GraphIndex uses the dreamdb.fresh-vamana-cosine discriminant and the versioned parameter shape below, because page reuse cannot be represented by the batch GraphIndex's current-snapshot hash binding.

4.1.1 Fresh GraphIndex parameter shape

The params map for dreamdb.fresh-vamana-cosine is CLOSED and has exactly these keys:

{
  "form": "fresh-vamana-params-v1",
  "base": {                         ;; exact spec/0013 §4.1 Vamana params
    "version": 1,
    "alpha": <f32-as-bytes>,
    "build_seed": <32 bytes>,
    "L_build": <uint>,
    "L_search": <uint>,
    "build_passes": <uint>,
  },
  "lineage_root": <multihash-GraphIndex>,
  "consolidation_threshold_appends": <unsigned int>,
  "consolidation_threshold_seconds": <unsigned int>,
}

The nested base map is CLOSED with exactly the six keys shown. Extension is by a new form, never by adding keys to v1. lineage_root is a reference that collectors and semantic replicators MUST traverse. An implementation that does not recognise the algorithm or form MUST refuse before querying, updating, or collecting the affected Manifest; the field is not an ignorable annotation (0002 §3.1.0).

Every Fresh snapshot names one existing batch dreamdb.vamana-cosine GraphIndex as lineage_root. Later snapshots keep it byte-identical. The current Track is the snapshot membership authority; the root is a stable graph-family identity, not a claim that every root page is still current.

4.2 Append semantics

For each new item v:

  1. Identify the entry point's nearest few neighbors via greedy search (per spec/0013 §4.4).
  2. The visited set becomes v's adjacency candidate set.
  3. α-prune to R out-edges (spec/0013 §4.3.3).
  4. For each accepted neighbor u of v: append v to u's adjacency list, then α-prune u's adjacency if it exceeds R. Collect old targets displaced by accepted reciprocal edges. Sort/deduplicate these targets by node id and require them in v's adjacency, filling any remaining slots from v's original pruned order. At most one target per reciprocal neighbor is displaced, so this still fits R. An old u → w path replaced by u → v remains reachable through v → w. If all reciprocal edges to v were pruned, v would be a published but unreachable node. In that case insert v into the entry point's adjacency. If that list was full, remove its last edge to w and ensure v's adjacency contains w (replace its last edge if full). Thus entry → w becomes entry → v → w without exceeding R. This deterministic admission bridge preserves that displaced route; it does not promise global ANN recall.
  5. Emit new GraphPage Objects for v and for any u that was modified. Their graph_index_hash is the lineage root, not the not-yet-known hash of the new snapshot. Reuse every unaffected page byte-for-byte.
  6. Publish a new Fresh GraphIndex Object with updated node_count and entry_point (the latter may shift slightly), the unchanged lineage_root. Per spec/0013 §3.1, GraphIndex is immutable; "updating fields" means emitting a fresh content-addressed Object at a new hash, not mutating bytes in place.

The new GraphIndex hash differs from the old; both remain content-addressed and immutable. The OLD Track Object remains valid for queries; the NEW Track Object becomes the latest after Manifest publish.

4.2.1 Snapshot and lineage validation

A reader MUST validate all of the following before assembling a Fresh snapshot:

  1. Fetch lineage_root, verify the Object against its address, and require a batch dreamdb.vamana-cosine GraphIndex. A missing root or another algorithm is a Protocol error.
  2. The graph-family fields match that root: dim, metric, R, page_node_count, page_bytes_target, vector_layout, and all six base Vamana parameters. Only entry_point, node_count, and the consolidation thresholds may differ.
  3. A lineage-preserving writer never reassigns an existing node id. An operation that changes a family field or renumbers nodes MUST build a new batch root and MUST NOT reuse pages from the prior lineage.
  4. node_count is at least the root's node_count. The current Track names exactly ceil(node_count / page_node_count) pages in contiguous page_index order. Each page's first_node_id, node count, R, vector width, modality, and final-page boundary agree with the current GraphIndex.
  5. Every current page's graph_index_hash equals the validated lineage root. It need not equal the current Fresh GraphIndex hash.

These rules make local reuse representable without weakening the existing lineage check. The root deliberately does not retain every intermediate snapshot: such a chain would grow without bound and still could not prove that a writer chose the best adjacency graph. The rules prove that the published snapshot is structurally one graph family and not a mixture of pages from independent graphs.

The SDK validates root identity and family metadata on open, the exact page count before query, and each page's headers/records when fetched. An explicit audit_graph_index walks every page. Lazy traversal does not claim to have audited unread pages. The initial runtime supports raw cosine roots (batch params v1 or identified v2); compressed Fresh roots are explicitly refused rather than guessed from the compressed batch profiles.

Dataset::append_graph_nodes(field, samples) adopts an existing raw batch graph on its first call, preserving node ids; empty samples can enable Fresh mode. The new field-scoped modality appends .fresh. The first adoption may copy unchanged bytes under this new namespace; later snapshots issue no PUT for an unchanged page address. Finite, nonzero input vectors are normalized; anchors must be unique across the existing graph and the entire batch, including reserved tombstoned slots. A u64::MAX anchor is refused because Track coverage is half-open. Every refusal precedes publication and the Manifest/Ref transition uses the ordinary optimistic CAS.

The initial writer decodes the current graph in memory, applies incremental search/pruning, and emits only changed pages. It does not run a full graph build on append, but it does not claim bounded memory or bounded read I/O for the writer. Query traversal remains lazy/cached. Out-of-core incremental writers are an implementation optimization, not an asserted result here.

4.3 Background consolidation

Append-only updates degrade graph quality over time (entry point drifts, alpha-pruning becomes sub-optimal). Periodic consolidation rebuilds the affected subgraph:

  • Triggered by SDK or operator schedule (default: every 1M appends or every 24 hours).
  • Selects regions of the graph touched by recent appends.
  • Re-runs the standard spec/0013 build algorithm over those regions.
  • Emits changed GraphPage Objects, reuses unchanged pages under the validated lineage root, and publishes a new Fresh GraphIndex via Manifest.

Consolidation is non-blocking for query path — readers continue to use the old GraphIndex while consolidation runs. Only after the consolidation Manifest is published do readers switch (via ref freshness).

The initial Dataset::consolidate_graph is caller-driven. Its region consists of nodes appended since the family root and nodes deleted or adjacent to a deleted node. It deterministically re-prunes live candidates (including the neighbors beyond a deleted bridge), clears deleted adjacency, and removes incoming edges to deleted nodes. A deleted entry point moves to the lowest live node id. If all slots are deleted the existing entry point remains, and tombstone filtering returns no rows. Node ids, payload slots and tombstones remain reserved; this operation is not physical erasure. Unchanged pages retain their addresses; an identical snapshot is a no-op, not a Ref advance. Re-pruning can alter approximate recall, not the identity of live records.

4.4 Delete semantics (tombstone-based)

Although DreamDB's data plane is immutable, logical deletes of indexed items are supported via the existing Layer mechanism (spec/0008): a "tombstone Track" carries the deleted doc-ids; query results filter against it. The graph itself is NOT modified — deleted nodes remain in adjacency lists but are filtered out of result sets.

Tombstones remain authoritative during and after an id-preserving consolidation. They may be retired only by an operation that actually removes the deleted records and proves no retained active source can make them visible again.

A consolidation that physically removes nodes and therefore renumbers any surviving node MUST emit a new batch GraphIndex root and rewrite its pages. It cannot reuse pages from the old lineage. A consolidation that preserves node ids may remain in the lineage and reuse untouched pages.

4.5 Consolidation parameter defaults

The CLOSED parameter shape is §4.1.1. Its consolidation_threshold_appends default is 1,000,000 and its consolidation_threshold_seconds default is 86,400. They are operator policy and may change between snapshots without changing the graph family.

4.6 Modality string

embedding.f32.dim=768.graph.r=64.field=features.fresh

The .fresh suffix is the marker; r is lowercase as required by the modality grammar, and the SDK binding is field-scoped. Without .fresh, the modality is non-streaming (rebuild-only) Vamana per spec/0013. The algorithm discriminant and params form, not this suffix alone, select the new page-lineage semantics.

5. Index health signal

Drift detection is the missing operational signal. Without it, recall degrades silently as the data distribution moves past the trained centroids/codebooks.

5.1 Manifest registry extension

Per-modality registry entry gains an OPTIONAL index_health sub-Object:

"registry": {
  "embedding.f32.dim=768.bucketed.spatial_bits=18": {
    …,
    "index_health": {
      "last_trained_at":      <u64>,                  ;; Unix ns when SpatialIndex was trained
      "last_trained_doc_count": <u64>,                ;; corpus size at training time
      "current_doc_count":    <u64>,                  ;; latest-known corpus size
      "estimated_recall_at_10": <f32-as-bytes>,       ;; SDK or operator measurement
      "drift_metric":         <f32-as-bytes>,         ;; centroid-shift L2; 0 = no drift
      "recommended_action":   "none" | "schedule-retrain" | "retrain-now" | "rebuild-graph",
    }
  }
}

5.2 Operator policy

  • recommended_action = "schedule-retrain": drift > 5% from training distribution. Schedule re-training next maintenance window.
  • recommended_action = "retrain-now": drift > 15% OR recall < 0.85. Immediate re-training advised.
  • recommended_action = "rebuild-graph": graph-based index whose consolidation backlog exceeds the per-modality consolidation_threshold_appends. Trigger consolidation.

These are non-binding hints — operators choose response — but a SDK MAY surface them in logs / metrics so operators don't miss drift.

5.3 How drift_metric is computed

For IVF / IMI: average L2 distance from each item's vector to its assigned centroid, vs the same metric at training time. A 10% increase in this metric ⇒ drift_metric = 0.10.

For LSH: compare the populated-cell distribution with a baseline measured for the exact 0004 §5.3 generator and representative source data. A uniform-on-sphere expectation is invalid for the published normalized-cube generator; departure from the versioned empirical baseline indicates drift.

For Vamana: average path length for greedy search from entry point to query target. A 50% increase ⇒ drift_metric scaled accordingly.

SDKs SHOULD compute drift on a sampling basis (1% of queries) and publish to the Manifest registry asynchronously via a Layer Track (per spec/0008). Drift estimation is NOT in the query hot path.

6. Re-training as a Manifest Layer

When operator policy triggers re-training:

  1. The training process produces a new SpatialIndex/VectorCompressor/GraphIndex Object.
  2. The operator publishes a new Layer Track (spec/0008) whose modality registry points at the new index.
  3. The old Track + old SpatialIndex remain reachable from the prior Manifest's parent chain; queries against old Manifests continue to work.
  4. New queries (via the latest ref) use the new index.

Re-training is therefore just another DreamDB publish operation — no special "migration" verb required. The cost is the wall-clock training time + the bytes of the new index. Queries during the training window use the old index (correct but slightly stale).

For full-corpus re-encoding (e.g., the corpus's vectors need to be re-quantized against a new codebook), see spec/0017 for the Reencode verb.

7. Conformance categories (per spec/0009 §8.6.2)

CategoryPass criterionCoverage
hotshard.append.flush-threshold.*Buffered items exceeding threshold trigger Track rewriteBoundary cases
hotshard.append.ttl.*TTL expiry forces flush even below thresholdClock-skew injection
hotshard.read.merge.*Query merges Track + HotShard candidates correctlyTime / vector / scalar predicates
hotshard.staleness.consumer-tolerance.*max_staleness_seconds honored; refresh triggered on missAll staleness levels
fresh-vamana.append-search.*After N appends, recall@10 stays within 5% of full-rebuild baselineN ∈ {10K, 100K, 1M}
fresh-vamana.consolidation.recall.*Post-consolidation recall returns to full-rebuild baselineAfter 10× threshold of appends
fresh-vamana.tombstone.delete.*Tombstoned items absent from query results; remain in graph until consolidationMixed insert/delete
fresh-vamana.page-lineage.*Unchanged root-bound pages remain reusable while rewritten pages and the current GraphIndex validate as one lineageRoot + two Fresh snapshots; foreign-root and missing-root negatives
index-health.drift-metric.*Drift metric monotonically tracks distribution shiftSynthetic drift injection
index-health.recommended-action.*Threshold transitions produce correct recommended_action transitionsAll policy levels

8. Latency and cost at scale

8.1 HotShard overhead

Per append (small batch, <100 items, embedding modality):

  • 1 GET of current HotShard (~50 KB warm) + 1 PUT of new HotShard (~50 KB) + 1 ref CAS.
  • ~30–80 ms p50 on commodity backends.

vs full Manifest publish:

  • 1 PUT of new Track Object + 1 PUT of any new Bucket + 1 PUT of new Manifest + 1 ref CAS.
  • ~200–500 ms p50.

~5–10× speedup for high-rate ingest.

8.2 Streaming Vamana cost

The following are design targets for an out-of-core writer, not measured properties of the initial in-memory incremental writer described in §4.2.1.

Append a vector to a 1B-item graph:

  • ~150 GETs of GraphPage (warm cache after the first ~50) → ~10 ms.
  • α-prune in-memory → <1 ms.
  • ~64 PUTs of updated GraphPage Objects (neighbors of v) → ~30 ms.
  • Total: ~40–60 ms per append. Throughput: ~20 appends/sec/connection.

At higher throughput, operators batch appends (each append's neighbor-updates can be merged) — typical workloads sustain 200+ appends/sec via batching.

8.3 Consolidation cost

1B-item graph; consolidation_threshold_appends = 1M (0.1% of corpus):

  • Affected nodes: ~10× the appends, so ~10M nodes.
  • Re-build cost: O(N × R × L_build) operations; ~30 minutes on a single multi-core machine.
  • New GraphPage emissions: ~10M nodes / 256 per page = ~40K pages = ~2.5 GB.

An operator may run consolidation in its own job; DreamDB does not start a background scheduler. Readers remain pinned to their snapshot until refresh.

8.4 Explicit maintenance planning (OQ-68)

The default is operator-controlled, with no automatic writes from planning. Rust Dataset::plan_maintenance(request) produces an ephemeral plan naming the immutable Manifest, Ref, exact scoped bindings, operation, limits and reasons. execute_maintenance(plan) performs at most one existing operation. It does not chain operations or publish through a second mechanism:

  • FlushHot requests a drain of all supported buffered fields. Planning reads HotShard Objects and reports their field-item count. Nonempty buffers qualify even when ingest is idle; the reason is an explicit drain, not an inferred TTL expiration. Empty buffers produce a no-op.
  • Compact selects one bucketed embedding field, a positive debt threshold and positive per-run cell rewrite limit. Planning explicitly uses 0021 §7.1's Exact scan, not metadata F as a proxy for debt. Oversized inputs qualify independently of optimization debt. The rewrite limit does not bound planning I/O. Execution delegates to ordinary compaction and its same global run limit.
  • ConsolidateGraph is an explicit operator request for an existing raw Vamana lineage. Quality/drift, append-since-consolidation, and elapsed-since- consolidation measurements are unavailable, not fabricated recommendations. It delegates to consolidate_graph, not bucket compaction or hot flushing.

Execution MUST refuse a different handle snapshot, Ref, or backend tip before writing. Each existing operation rechecks its parent and retains its normal CAS protection against subsequent races; a losing CAS may leave unreferenced content Objects, exactly as for direct invocation. Plans cannot be silently refreshed or retargeted. Re-plan after a stale refusal. No-op plans also check staleness.

No resident daemon, automatic mode, registry scheduler key, retraining policy, consumer-refresh SLA or new durability exception is defined. HotShard TTL checks inside append_hot are not a timer: stopping writes does not itself flush.

9. Out of scope

  • Linearizable cross-writer freshness. HotShard provides per-writer monotonic freshness; cross-writer requires consensus, which DreamDB defers (spec/0008).
  • Cross-modality drift detection. Per-modality only; correlated drift across multi-modal tracks is application-layer.
  • Adaptive index reconfiguration. "Switch from IVF to IMI as data grows" is a future operator tool; spec doesn't automate it.
  • Online learning of codebooks. QINCo codebooks are fixed at training time; online updates would change the content-hash, breaking immutability. Re-training is via Layer publish (§6).

10. Open questions

  • OQ-67 (→ this spec): Whether a future explicitly distinct hot-tier mode may weaken durability. The existing published HotShard is ordinary committed content and remains subject to the backend contract; this question is not permission to acknowledge an undurable v1 commit. Filesystem fsync and remote PUT acknowledgements require their own backend guarantees.
  • OQ-68: Resolved. §8.4 chooses explicit operator planning/execution and separates the three existing operations. Automatic scheduling is deferred, not shipped.
  • OQ-69 (→ this spec): Drift metric definition for SPLADE / ColBERT modalities. Centroid-distance doesn't generalize cleanly. Likely a different per-algorithm metric. Defer until SPLADE/ColBERT implementations land.
  • OQ-70 (→ spec/0009): Portable vectors for streaming-Vamana correctness, including incremental-build determinism. Required before claiming that capability's conformance; not a release gate for the existing HotShard path or static Vamana.

Next: spec/0017 — schema evolution. When the embedding model upgrades, we need to re-index 10B items without losing the old corpus. Bridges the gap from spec/0016 incremental updates to "bulk migration."