Spec 0017 — Schema Evolution and Embedding Migration
Status: Draft (Phase 4 design).
Depends on: spec/0001, spec/0002, spec/0006, spec/0008, spec/0010, spec/0016.
Motivation: A 10B-item DreamDB deployment outlives any one embedding model. OpenAI's text-embedding-3 obsoleted text-embedding-ada-002 within 14 months; clip-ViT-B/32 has been the dominant image encoder for 3 years and will eventually be replaced. When the model upgrades, the operator faces a brutal choice today: keep the old corpus and accept degraded quality on new queries, or re-encode 10B items in one atomic operation that takes weeks. Both are wrong. spec/0017 defines the protocol-level primitives — multi-version modality registries, a Reencode verb, and a compatible_with hint — that make incremental, partial, resumable migration the default.
1. Purpose
The protocol's immutability and content-addressing make schema evolution structurally easy: a new modality is just a new modality, with its own Track Object, its own SpatialIndex, its own VectorCompressor. The hard part is migrating gracefully — keeping queries working through the transition, sharing storage where possible, and avoiding the "stop the world" rebuild.
By the end of this document the following are concrete:
- Multi-version modalities in a single Manifest's registry —
embedding.v1andembedding.v2register independently, share Items, and queries route to the right version per-call. - The
Reencodeverb: operator-driven bulk re-index that reads source Items, applies a transform (typically: run new model inference), writes target Items, and incrementally publishes progress. Resumable, idempotent. - The
compatible_withregistry hint: optional declaration that a new modality is approximately compatible with an old one — used by the query planner (spec/0015) for explicitly requested multi-version result fusion. Per-Item migration fallback uses §7'sBackfillClaim, never this hint. - Versioned modality strings: a discipline for naming evolving modalities so old/new tracks coexist without ambiguity.
- The migration manifest pattern: how a long-running re-encode produces a sequence of Layer Manifests rather than one giant atomic commit.
What stays defined elsewhere:
- Per-modality storage layouts — spec/0007, spec/0010, spec/0013.
- The Layer mechanism — spec/0008 §3.
- Streaming updates / hot-shard — spec/0016.
What this document does NOT define:
- Automatic model selection. Which model to migrate TO is operator-policy.
- Transform correctness verification. That the new model's outputs are "right" is the operator's training-eval concern, not the protocol's.
- Cross-modality lossy conversion. Converting a CLIP embedding to a BERT embedding is meaningless and out of scope.
- Model deployment. How operators run inference at scale is implementation-defined; this spec defines only the DreamDB-side coordination.
2. Multi-version modalities
2.1 Modality versioning convention
A version-aware modality string carries an explicit version=<N> parameter:
The version parameter is OPTIONAL. Without it, modalities are unversioned (effectively version=1 implicit). Two modalities with the same shape but different version are distinct modalities for all protocol purposes — different path slots, different Track Objects, different SpatialIndex Objects, different bucket headers.
This is intentional. Modality strings are content-addressing keys; two modalities are "the same" iff their strings are identical. The version parameter makes incompatibilities visible.
2.2 Multiple versions in one Manifest
A Manifest registry MAY declare multiple versions of the same logical concept:
Both versions are individually queryable; the second's compatible_with field declares its relationship to the first. The query planner (spec/0015) uses this to route hybrid queries across the migration boundary.
compatible_with.coverage is a planner hint and is not backfill coverage.
It may inform whether an application offers multi-version fusion, but it never
selects individual fallback Items. Section 4.2 uses only a BackfillClaim for
that decision.
It is not normative backfill coverage (§7) and is never a fallback for a missing BackfillClaim: an absent claim means no backfill contract, not "consult the hint". No value is projected between the two in either direction; a claim is never derived from compatible_with, and compatible_with.coverage is never derived from a claim.
The reason is structural rather than stylistic. compatible_with.modality names a type, with no address and no basis, so the relationship cannot identify which Items a completeness statement is about. §7 needs exactly that, which is why it carries its own basis.
Arrays remain forbidden from using compatible_with at all (0025 §8); §7 costs them nothing, because backfill coverage does not travel this path for anyone. Removing the planner hint, if it is ever removed, is separate work.
2.3 Coverage during migration
During a long-running migration, coverage = "partial" indicates that not every v1 Item has been re-encoded to v2 yet. The planner's behavior depends on the query:
- Query against v2 explicitly: returns only v2 results. Coverage gap is visible to the application.
- Query against v1 explicitly: returns v1 results (unchanged).
- Query against the logical concept with
version_preference = "all": planner queries both versions and performs explicit multi-version result-set fusion. This is not coverage fallback. - Query against the logical concept with
version_preference = "coverage-fallback": planner uses the target binding's §7BackfillClaimand its exact basis to fill only Items whose result isUndecided.
The "logical concept" path requires a small extension to the query verb's track_selector — see §4.
The coverage value read here is the §2.2 planner hint. It is a routing signal, not a per-Item statement, and it cannot answer "does this Item have a value" — that is §7's question, and §7's claims answer it against an explicit basis.
3. The Reencode verb
A new verb. Reencode reads source Items from a source modality, applies a transform (operator-supplied function), writes target Items to a target modality, and publishes progress as Layer Manifests.
3.1 Verb signature
The RPC sketch's capability bytes have no portable DreamDB token encoding.
A deployment exposing this RPC must authenticate under its own profile and
authorize source reads and target writes (0018 §3); accepting opaque bytes is
not verification. Credentials are runtime inputs, not persisted Reencode state.
The SDK implementation walks the source Track's Items in anchor order, applies the transform (out-of-band — DreamDB doesn't dictate how), and appends to the target Track. Progress is checkpointed every batch_size Items via a Layer Manifest.
3.2 What is transform_ref?
An opaque content hash. The protocol does NOT define what bytes it points to — different deployments use it differently:
- Inference deployments:
transform_refpoints at a small CBOR Object describing model identity, weights hash, preprocessing config. The SDK uses this to look up the correct inference endpoint. - Pure-transform deployments (e.g., re-normalizing existing vectors):
transform_refpoints at a CBOR Object describing the math. - Test deployments:
transform_refis a no-op identifier; the SDK skips actual encoding.
DreamDB stores the hash as audit trail. The operator's external system resolves it.
3.3 Idempotency and resume
Reencode publishes intermediate progress as Layer Manifests, each one valid as a queryable state. A crash mid-Reencode leaves the system in a consistent partial state — the next invocation with resume_from: <last-published-Manifest-hash> picks up at the next batch.
The Layer Manifest's body carries a small reencode_state sub-Object:
Resume reads this; verifies the source modality has not changed since the checkpoint (Manifest parent-chain walk); continues from last_anchor + 1.
3.4 Concurrency with live ingest
Reencode runs in parallel with live writers appending to the source Track. The migration's coverage view treats anything in the source's HotShard or appended after last_anchor as "not yet migrated"; new appends after checkpoint_at are flagged as backlog for the next Reencode pass.
For a continuously growing Track, full coverage = "all items written before the final Reencode pass." Operators typically run several passes:
- Pass 1: covers the bulk corpus at time T1. Coverage at completion: items with anchor < T1.
- Pass 2: covers the backlog (items added between T1 and T2). Faster, smaller.
- Pass 3+: convergence as backlog shrinks.
Eventually the operator declares "migration complete" and updates the coverage field to "complete" (or removes the v1 modality from active registry).
This watermark model is Reencode's and is not backfill coverage. A source-derived re-encode uses last_anchor as its ordered-progress marker. §7's claims instead quantify over a frozen cohort: decisions may be made out of ordinal order but their set grows monotonically across claim succession. A later append outside that cohort can provide a readable Value; it never enlarges the basis or contributes to the claim's decided set. A missing value outside the basis is OutsideDomain (§7.11), not a backfill decision.
3.5 Resource budgeting
Reencode is resource-intensive (one inference forward pass per Item × 10B Items can take days). The verb body MAY include budgets:
Budgets govern admission of transform batches, not preemption of an admitted
batch. Before invoking each transform the driver MUST check deadline_at;
if it has arrived, no new batch is started. A batch already admitted completes
its ordinary atomic data-and-checkpoint publication even if the deadline passes.
This is not a hard wall-clock deadline on inference, I/O, source enumeration or
publication. Source-window collection remains unbounded by these controls.
max_items_per_hour, when present, MUST be positive. Each invocation allows
one initial batch as a burst; subsequent admission waits until elapsed monotonic
time is at least ceil(admitted_source_items * 3600000000000 / rate) nanoseconds.
Count source items handed to the transform, including a batch yielding no output.
The rate accounting is per invocation, not a persistent or cross-client quota.
Restarting/resuming grants a new initial burst. No wait is needed after the last
batch. If the next permit cannot arrive before the deadline, return the existing
progress immediately; recheck the deadline after waiting before admission.
A pause with no newly published batch MUST NOT synthesize a checkpoint or move
the Ref. batches_published is zero and final_manifest names the unchanged tip.
On an initial pass that tip need not contain a migration checkpoint: restart with
no resume_from. On a resumed pass retain the original resume_from if nothing
was published. Once a batch is published, its final_manifest is the resume point.
The reference native implementation enforces this admission policy. A runtime
without a rate timer (currently wasm32) MUST reject a requested rate before
transform or writes, never silently ignore it. The reference transform is serial;
the former illustrative max_concurrent key is not an implemented budget and
is not part of this contract. These controls do not impose memory, accelerator,
tenant-wide or backend quotas. Those require separately specified mechanisms.
Compatibility: previously an expired deadline still admitted the first batch, zero rate acted unlimited, and wasm32 skipped rate sleeps. Those behaviors are not the admission contract; callers wanting unlimited work omit the budgets.
4. Query planner extensions (spec/0015 amendment)
4.1 Logical track selector
The HybridQuery's track_selector (per spec/0015 §5.1) gains an OPTIONAL logical_concept field:
logical_concept is a free-text label; version_preference controls planner behavior:
"latest": query the highest-version registered modality of this concept. A coverage gap is not filled."all": query every version and merge their result sets via the planner's hybrid fusion. Membership in a top-K result is a ranking fact, not a coverage fact; this policy makes no coverage claim."coverage-fallback": query the latest target and the exact basis named by itsBackfillClaim, then fill onlyUndecidedbasis Items under §4.2. A missing claim is an error, not an invitation to infer coverage fromcompatible_with."<version-spec>": query a specific version (e.g.,"version=2").
The default is "latest" — applications get the new model's results unless
they explicitly opt into another policy. Migration tooling that wants to fill
an incomplete backfill SHOULD select "coverage-fallback"; applications that
want independent multi-model fusion select "all".
4.2 Coverage-fallback logic
For "coverage-fallback", the target binding MUST carry a valid §7
BackfillClaim. The fallback source is that claim's exact basis binding,
not whichever older modality happens to return a result and not a type named
by compatible_with:
The older-version penalty is a small but non-zero discount that prefers target embeddings over genuinely pending basis values. Default 0.9; operator-tunable. Implementations MAY batch and cache claim resolution, but MUST preserve the four-state decisions above.
Two cases distinguish this policy from result-set fusion:
- Item X has a target
Value, ranks below the target sub-query's returned pool, and ranks highly in the basis sub-query.coverage-fallbackexcludes X's basis score because X is covered;allmay include X through its explicitly requested fusion policy. - Item Y is
Undecidedin the target claim and ranks highly in the basis sub-query.coverage-fallbackincludes Y's penalized basis score because Y is a real coverage gap.
Testing membership in the target result set cannot distinguish X from Y and MUST NOT be used as a substitute for resolving the claim.
4.3 Score scale calibration
Different model versions produce different cosine-similarity distributions. Linear fusion across versions risks miscalibration. RRF (spec/0015 §5.2) is scale-invariant and is the recommended default for multi-version hybrid queries.
5. Garbage collection across versions
The protocol's GC (spec/0006 §7.3) is purely content-reachability — Objects unreachable from any live Ref are eligible for deletion. Multi-version registries simply keep the old version's Track + SpatialIndex reachable until the operator explicitly removes them from registry.
A backfill claim adds one reachability obligation, and it is not conditional on the claim being incomplete: while any readable claim exists, its basis and its decided-set Object remain reachable regardless of state. Retiring the basis from active into lineage is always permitted; dropping it from both is not, until no reachable claim references it. See §7.8.
5.1 Decommissioning v1
When the operator decides v1 is no longer needed:
- Publish a Manifest whose registry omits v1 entirely.
- Old Manifests in the parent chain still reference v1's Track Object — they remain reachable until the parent chain is GC'd.
- Eventually, after the GC's safety threshold (default 24h, spec/0006 §7.3), v1's Objects become eligible for deletion.
- Optional: a "snapshot roll-up" (spec/0008 §9.3) accelerates GC by collapsing the parent chain.
5.2 Concurrent migration safety
Two operators running parallel migrations to different target modalities (e.g., v2 and v3 concurrently) is supported — each Reencode publishes its own Layer Manifest. The standard spec/0008 merge / rebase rules apply at Publish time.
6. Worked example: 10B CLIP-B/32 → CLIP-L/14 migration
Concrete scenario. DreamDB deployment with 10B image embeddings under embedding.f32.dim=512.bucketed.spatial_bits=22.version=1 (CLIP-B/32). Operator wants to migrate to CLIP-L/14 (dim=768).
Step 1: register v2 alongside v1.
Manifest registry now has both. coverage = "partial" for v2; items_done = 0.
Step 2: invoke Reencode.
Step 3: Reencode publishes Layer Manifests every 1M Items. After ~100 hours (at the budgeted rate), all 10B Items re-encoded.
Step 4: the operator completes v2's §7 claim. Queries via
version_preference = "coverage-fallback" now have no Undecided basis Items
to fill; version_preference = "all" remains an explicit multi-version fusion
request and does not silently change meaning.
Step 5 (optional): operator omits v1 from a subsequent Manifest's registry. v1's Objects become GC-eligible after the safety threshold.
Cost:
- Inference: 10B forward passes × ~5 ms = ~14000 GPU-hours (operator-side, out of band).
- Storage: ~10B × 3 KB (v2 uncompressed) = 30 TB during transition (v1 + v2 both present). After GC: 30 TB (v2 only; v1 freed).
- Network: ~30 TB outbound from compute layer to backend; standard.
- Wall clock: bounded by inference, not DreamDB.
The DreamDB side adds zero new failure modes — partial state is always a valid queryable Manifest; resume is idempotent; rollback is just "publish a Manifest that omits v2."
7. Backfill coverage
Resolves OQ-102.
7.1 The problem a claim solves
An optional field that a writer does not set produces no record in that field's Track. A field that has not been backfilled yet also produces no record. On the wire the two are identical, so a reader cannot distinguish this Item has no value from this Item has not been decided yet — and a reader that guesses picks silently, which is the shape that produces a corrupt training set from a healthy-looking dataset.
Nothing can recover that distinction from the data. It requires a positive declaration: a statement, written by the party that ran the backfill, of which Items it undertook to decide. That declaration is a backfill claim.
0002 §7.2.6 defines the claim's bytes and the container's local rules. This section defines what a claim means.
7.2 Target and basis
A claim names two binding versions, and conflating them is the error this design exists to avoid:
| Question it answers | Value | |
|---|---|---|
target | Which binding version is this claim about? | the exact BindingVersionRef being backfilled |
basis | Which Items does it undertake to decide? | an immutable binding whose Items are the claim's domain |
A modality is type identity and may be shared by many fields and bindings; a BindingVersionRef is (timeline, field) plus modality plus the Track's content address, so it is instance identity. Coverage is instance state, and MUST NOT be keyed by modality.
In v1 the basis is an explicitly chosen sibling BindingVersionRef, and nothing else. It MUST:
- lie on the same Timeline as the target;
- be immutable — which it is, being pinned by its Track content address;
- have unique anchors (§7.2.1);
- be named explicitly by the writer. It is never inferred and never defaulted.
The set that binding enumerates is the declared domain. A claim is never a statement about the Dataset as a whole, and MUST NOT be read as one. If the chosen sibling's Track omits Items, those Items are outside the domain and the claim says nothing about them — which is why the choice must be explicit rather than derived.
Where no suitable sibling exists, v1 does not support that backfill: publication is refused rather than approximated. A dedicated basis snapshot Object is deliberately not introduced.
For an item-type change (0025 §8) the basis is the predecessor binding, which already satisfies every requirement above.
7.2.1 Unique anchors
Publishing or resolving a claim MUST be refused when its basis contains duplicate anchors.
This restriction binds only a claim's basis. It is not a statement about Tracks in general, it does not apply to unrelated Tracks, and it does not retroactively invalidate stored data. 0021 §3.2 permits duplicate anchors within a Track — identical record bytes are deduplicated and only differing bytes fail — so the corpus tolerates elsewhere what a basis refuses.
It exists because a bare anchor cannot address two Items that 0001 §5.4 explicitly permits to share a timestamp:
two distinct samples can share a timestamp
0001 §5.4 describes application-level identity lookup through a scalar id field indexed via 0011; that lookup does not enforce uniqueness or change this basis's anchor-uniqueness requirement. The optional typed entity-key contract in 0026 resolves OQ-88, but does not replace the frozen basis's anchor-based addressing defined here. OQ-102 does not redefine entity identity. A duplicate-anchor basis is refused, never guessed at.
Uniqueness is also what makes §7.3's ordinals well-defined.
7.3 The basis ordinal
Because the basis is immutable and anchor-unique, its Items have a canonical total order: ascending anchor. The ordinal of a basis Item is its 0-based position in that order. N is the basis Item count.
N and all ordinal endpoints are u64.
Ordinals are dense over [0, N) by construction, and this is the identity space in which every extent, decision, succession comparison and merge in this section is expressed. 0002 §7.5.1 gives the reason extents are ordinals rather than anchors: anchor extents are not canonical over a sparse basis.
7.4 The basis is a frozen cohort
A claim describes a fixed historical basis. Items appended after the claim was made do not join it.
The basis is immutable because nothing is ever added to it, and complete is verifiable because it quantifies over a finite, pinned set.
Two consequences, stated here as limits rather than left to be discovered:
- §3.4's
last_anchorwatermark is not this. It remains correct for Reencode, which reads Items that already exist in anchor order, and is not changed by this section. It does not define backfill coverage: a claim quantifies over a frozen cohort, not a moving frontier, so a claim's decided set need not be monotone in anchor order and that is not a defect in the watermark model. - An optional field omitted by a write outside every frozen basis is outside the field's declared domain. Such an Item is not in any cohort, so no positive statement makes the omission
Absent. Its public read result MUST beOutsideDomain(§7.5 row 6), and no API may map it toAbsent— not by default and not behind a flag. Folding it would recreate exactly the ambiguity §7.1 removes, one indirection further from the reader. §7.11 fixes this as the v1 ongoing-write rule rather than creating a second, moving coverage system.
7.5 Reading: a six-cell grid
Positions are basis ordinals. Every combination is defined.
| # | Basis membership | Extent | Target record | Result |
|---|---|---|---|---|
| 1 | ordinal i ∈ [0, N) | inside a decided range | present | Value |
| 2 | ordinal i ∈ [0, N) | inside a decided range | none | Absent — decided, and the decision was "no value" |
| 3 | ordinal i ∈ [0, N) | outside every range | none | Undecided |
| 4 | ordinal i ∈ [0, N) | outside every range | present | Protocol error (§7.9) |
| 5 | not in the basis | n/a | present | Value |
| 6 | not in the basis | n/a | none | OutsideDomain |
Rows 3 and 4 share an extent position and are separated only by whether a target record exists: row 3 requires none; row 4 is the present case.
Undecided is a normal, explicit read result — not an error and not a failure. A point read landing on row 3 returns it. What a reader MUST NOT do is report row 3 or row 6 as row 2.
Row 5 lies outside the claim. Such a Value is readable, but it does not contribute to complete and does not participate in decided-set validation. The coverage domain is not the readability domain: a claim narrows what can be concluded, never what can be read.
complete means every basis Item has been decided. It does not mean every Item has a value — row 2 is a decision — and it says nothing about Items outside the basis.
7.5.1 Several target records at one basis ordinal
The basis is anchor-unique (§7.2.1); the target Track is not. Two target records may carry the anchor of one basis Item.
- Identical bytes collapse to one logical
Value, consistent with0021§3.2's deduplication. - Differing bytes are a Protocol error. Two incompatible values for one decided Item MUST NOT be resolved by a first-wins or last-wins rule.
This applies at rows 1 and 5 alike. At row 4 the undecided-position refusal takes precedence: a record that should not exist is not improved by there being two of it.
7.6 state, validated against the decided set
state restates facts already carried by the decided set, so it MUST agree with it in both directions. Let D be the decided ordinal set and N the basis Item count.
Empty basis (N = 0): the only legal state is complete.
An empty cohort has nothing left to decide, so complete is the honest reading, and it is also the only one that keeps the encoding unique — under the general rules below an empty basis would satisfy not-started (D = ∅) and complete (every basis Item decided, vacuously) at once. not-started and partial are refused when N = 0.
Non-empty basis (N > 0):
not-startediffD = ∅partialiffD ≠ ∅and|D| < Ncompleteiff|D| = N
Any other pairing is a Protocol error. There is no distinct meaning for an empty partial, so it is not permitted: not-started has exactly one encoding.
Every extent endpoint MUST satisfy end ≤ N. 0002 §7.5.1 defines the Object's own normalisation; this bound needs the basis and is checked here.
7.7 Claim lifecycle
7.7.1 Genesis
An initial claim — one with no parent claim for the same target logical binding — MAY declare not-started, partial or complete, provided every local rule holds. A writer is not required to publish an empty claim first and then progress it; requiring that would add a publication that proves nothing.
An empty basis remains complete only (§7.6).
A Manifest with parents: [] is a legitimate claim genesis, not a missing-history failure. Per 0008 §9.3 a snapshot roll-up publishes exactly such a Manifest and is explicitly a new, lossy trust root: history before it is gone by design. Only the local claim rules can be proved across that boundary; succession (§7.7.3) cannot, and a resolver MUST NOT report an unprovable succession as a violation when the boundary is a legitimate root.
7.7.2 Continuity
Once a live claim (0002 §7.2.6) exists for a reachable target, a child Manifest MUST NOT silently drop it while that target's logical binding remains in active or lineage. A claim disappears only with its target: when the target leaves both active and lineage, its claim is removed with it.
Dropping a claim while keeping the binding would return the field to the §7.1 ambiguity without saying so — the one outcome this section exists to prevent.
7.7.3 Succession
Every backfill batch that writes a new target Track produces a new BindingVersionRef. But progress does not require one. A batch may decide Items only as Absent, or bring an already-present Value inside the decided set; it writes no target record, so the Track bytes and therefore the BindingVersionRef are unchanged and only decided_ref moves.
A claim C' succeeds C iff all hold:
C'.targetis either exactlyC.target, or an ordinary successorBindingVersionRefwith the same logical binding(timeline, field)and the samemodality;C'.basisis identical toC.basis. Changing the basis is not succession and MUST be refused — it is a different domain wearing the same name;C'.decided ⊇ C.decidedas ordinal sets. The decided set may be retained or expanded and MUST NOT shrink;- every ordinal
Cdecided is decided the same way byC': a row-1ValuestaysValue, a row-2AbsentstaysAbsent. Re-deciding an already-decided Item to the other outcome is refused.
A pure append outside the basis may carry the claim forward with an unchanged decided set: the new Item is not in the frozen cohort, so nothing about the cohort changed.
The old target remains reachable wherever ordinary lineage rules require it (§7.8).
7.7.4 Ancestry: linear and merge
- In a linear history, the new claim succeeds the relevant direct-parent claim.
- In a merge, the new claim MUST preserve and supersede every relevant direct-parent claim. Parent claims for the same target logical binding that name different bases are a conflict, not merge inputs.
7.7.5 Who validates what
| Layer | Obligation |
|---|---|
Manifest decoder (0002 §7.2.6) | every local rule: CLOSED schema, ordering, claim uniqueness, target/basis resolution and Timeline agreement, liveness, and the decided-set Object's own normalisation. All decidable from one Manifest |
| Dataset resolver — whatever component consumes coverage to answer a read | succession and continuity: recursively verify the reachable parent DAG, checking monotonicity and decision preservation |
The resolver MAY cache verified ancestry. A named parent that cannot be fetched is a Protocol error — an unreachable parent is not an absent one. A Manifest that legitimately declares parents: [] (§7.7.1) is genesis and terminates the walk.
Writers MUST perform these checks before publishing; a Manifest violating succession is a malformed write. The resolver MUST NOT rely on that, for the same reason 0002 §7.2.3 has readers re-apply the size limit independently: a hand-built Manifest must not bypass a writer. The parent-chain walk is the pattern §3.3 already uses for Reencode resume.
7.8 Reachability for the lifetime of the claim
A complete claim still needs its basis — to map an anchor to an ordinal (§7.3), to separate row 2 from row 6, and to verify complete at all. Reachability is therefore not conditional on the claim being incomplete.
- While any readable claim exists, its
basisand its decided-set Object remain closure-reachable, regardless ofstate. - Moving the basis from
activetolineageis always permitted when query fallback does not need it. Retirement from the query-visible set does not break a claim, which addresses the basis byBindingVersionRef, and garbage collectors traverselineage(0002§7.2.3). - Final garbage collection is permitted only after no reachable claim references the basis — again regardless of
state.
Query-visibility, provenance retention and collectability are three separate transitions, and 0002 §7.2.3 already separates them; this rule is written in those terms rather than adding a fourth.
7.9 Failure behaviour
The following are Protocol errors — published metadata contradicting published Items, or contradicting itself. The Schema is coherent and the caller supplied nothing wrong; what failed is the object.
- §7.5 row 4 — a target record at an undecided ordinal.
- §7.5.1 — differing target bytes at one basis ordinal.
- §7.6 — any
state/ decided-set pairing that breaks the biconditional, including a non-completeempty basis. basisunresolvable, or resolving to a basis with duplicate anchors (§7.2.1).decided_refunresolvable, extents not normalised (0002§7.5.1), orend > N.- §7.7.3 — a succession violation found on the parent DAG.
- §7.7.5 — a named parent that cannot be fetched.
Fail closed, never repair on read. A reader that narrows the decided set to fit the records it happens to find has invented a claim nobody wrote.
7.10 No claim is not a state
not-started is available only once a claim with a non-empty basis exists, and means exactly "this basis is declared, none of it decided".
A binding with no claim at all is not not-started. It has no basis, therefore no domain, therefore nothing to be complete about, and every read falls in §7.5 row 5 or row 6. Reporting that as not-started would assert an undertaking nobody made.
7.11 Ongoing writes outside a frozen basis
For v1, an ongoing write does not create an implicit coverage claim. The meaning is determined by facts actually published:
- a target record outside every claim returns Value (row 5);
- no target record outside every claim returns OutsideDomain (row 6);
required: truestill makes omission a writer-side Schema error; andrequired: falsepermits omission but does not turn it into a positiveAbsentdecision.
This is a deliberate protocol boundary, not an unknown answer. A caller that needs a durable present/absent distinction for every newly appended Item must make that distinction data: use a required field, or write a real value in the declared item type whose application semantics express the sentinel. DreamDB does not invent a per-Item marker or a rolling basis on the caller's behalf.
The alternatives rejected for v1 each add state without an existing identity model that can validate it:
- a per-write absence marker adds per-Item metadata to the payload model that
0025§11 deliberately keeps out; - a mutable or rolling basis makes
completedepend on a moving cohort and breaks the frozen-basis invariant in §7.4; and - interpreting every optional omission as
Absentwould retroactively change the meaning of releasedrequired: falsedata whose bytes carry no such statement.
A later protocol may add an explicit positive-decision mechanism under a new wire form. It MUST be opt-in and MUST NOT reinterpret an omission published under this v1 rule.
8. Conformance categories (per spec/0009 §8.6.3)
| Category | Pass criterion | Coverage |
|---|---|---|
evolve.multi-version-registry.* | Registry with v1 + v2 both query correctly | Both versions, independent queries |
evolve.reencode.resumable.* | Crash mid-Reencode + resume produces same final state as uninterrupted run | Failure injected per-batch |
evolve.reencode.checkpoint-monotonic.* | last_anchor strictly increases across batches | Adversarial batch orderings |
evolve.planner.all-versions-fusion.* | version_preference: "all" fuses independently returned version result sets without claiming coverage semantics | Covered-but-outside-target-top-K case |
evolve.planner.coverage-fallback.* | Only Undecided basis Items receive fallback scores; Value, Absent, and OutsideDomain do not | §7 outcomes including covered-but-outside-target-top-K |
evolve.gc.decommission-v1.* | After v1 removed from registry, its Objects become eligible after safety threshold | Standard GC test |
evolve.compatible-with.semantics.* | coverage remains a coarse planner hint and never selects per-Item fallback | Missing claim plus Value / Undecided counterexamples |
backfill-read-grid | Each of §7.5's six cells is exercised, and the empty-basis state resolves | Value / Absent / Undecided / OutsideDomain / row-4 Protocol; N = 0 |
backfill-decided-set-roundtrip | 0002 §7.5.1 bytes round-trip, and a non-canonical ordinal set is rejected | Coalescing, ordering, overlap, start < end, v, two-element ranges |
backfill-basis-anchors | A basis carrying duplicate anchors is refused at publish and at resolve | §7.2.1 |
backfill-succession | Same-target Absent-only progress and new-target progress are accepted; shrink, decision flip, changed basis and unrelated target are refused | §7.7.3 over a parent DAG |
backfill-merge | Identical target bytes deduplicate; differing bytes and Value-vs-Absent conflict | §7.5.1, §7.7.4 |
The five backfill-* categories are fixture-driven per 0009 §3: a third-party implementation runs them from the vector inputs without linking these crates.
9. Out of scope
- Lossy embedding-space alignment. Migrating from one embedding model to another with substantially different geometry (e.g., 1024-dim → 1536-dim) is a research problem; DreamDB stores the bytes, not the alignment.
- Cross-model query. "Search v1 corpus using a v2 query" requires a learned alignment matrix; out.
- Schema diff / merge tools. Out of protocol; SDK / CLI concern.
- Automatic transform validation. That
transform_refactually corresponds to model X is the operator's audit problem.
10. Open questions
- OQ-71 (→ this spec): Should
compatible_withdeclare a similarity-space transform (e.g., a learned alignment matrix hash) for cross-version score combination? Currently we use the older-version penalty heuristic; a learned alignment could be more principled. Defer to v0.X+1. - OQ-72: Resolved. §3.5 defines enforced batch-admission rate/deadline controls, resumable pauses and explicit unsupported-runtime refusal (#356). Admitted batches finish atomically; this is not encoder preemption, a memory bound or a cross-client quota.
- OQ-73 (→ spec/0006): RESOLVED.
0006§2.2 defines Reencode as a distinct optional extension to the eight core verbs, not an eleventh core verb or an alias for Append. A claim to implement Reencode carries that section's resume obligations. - OQ-74 (→ spec/0009): Portable Reencode failure/resume vectors. Reference implementation tests are not independent-implementation evidence. This blocks a portable Reencode conformance claim, not unrelated capabilities or releases.
OQ-102 is resolved by §7. It asked how a reader distinguishes a genuinely absent value from one not yet backfilled; §7 answers it with an explicit claim over a frozen, anchor-unique basis, and 0002 §7.2.6 / §7.5.1 give the wire form.
§7.11 also resolves the ongoing-write question for v1: an optional field omitted outside every frozen basis is OutsideDomain, never Absent. No implicit claim, per-Item absence marker or rolling basis is created, and a future opt-in wire form must not reinterpret data written under this rule.
Next: spec/0018 — multi-tenant operation. Now that 10B-scale + federation + hybrid + streaming + schema evolution all work for ONE tenant, can they work for many at once without collapsing into noisy-neighbor chaos?