{/* Generated from dreamdb-core by scripts/sync-spec-docs.mjs — do not edit. */}
# Spec 0024 — Embedding Spec Identity

**Status:** Implemented.
**Depends on:** `spec/0001`, `spec/0002`, `spec/0004`, `spec/0015`, `spec/0017`.
**Motivation:** Two vectors of the same dimension are silently comparable. Nothing in the protocol prevents a query encoded by one function from being scored against an index built by another — cosine similarity computes fine, scores look plausible, and the results are confidently wrong. This has already happened in practice, and it is not detectable from the data. spec/0017 makes migration between embedding versions graceful; this document defines what makes two embeddings *the same version in the first place*, and requires that mismatches fail loudly.

---

## 1. Purpose

`spec/0017` defines versioned modality strings and the `compatible_with` hint so old and new tracks can coexist through a migration. It explicitly does **not** define what makes a vector belong to one version rather than another — it treats that as operator policy.

That gap is a correctness hole, not a policy question. This document closes it:

- **The `embed_spec`**: the complete, pinned description of the deterministic function that produced a vector.
- **The `spec_id`**: a content hash over the parts of that description which, if changed, produce an incomparable space.
- **Modality binding**: the `spec_id` travels in the modality string, so incompatible embeddings cannot occupy the same field.
- **Mandatory validation**: where the check happens and what it must refuse.

What this document does NOT define:

- **Which model to use.** Operator policy, as in spec/0017.
- **Whether a model is good.** Retrieval quality is an evaluation concern; see the acceptance-gate discipline in §6.
- **Storage layout.** Unchanged — spec/0007, spec/0010.

---

## 2. The problem is that dimension is not identity

A 768-d SigLIP vector and a 768-d CLIP ViT-L/14 vector are structurally indistinguishable. Every layer of the system accepts both:

- `Dataset::query_vector_scored` validates `query.len() == schema.dim` and nothing more.
- `IvfCosine::hash_vector` computes dot products against centroids regardless of provenance.
- The rerank pass computes exact cosine on f32 that means something different.

The failure is **silent, plausible, and total**. There is no error, the score distribution looks normal, and every returned record is wrong.

**Identity is the function, not the model.** Measured evidence, `google/siglip-base-patch16-224` text tower, 300 prompts, pooler-output cosine against the PyTorch source:

| export | parity (mean / min) | label-as-query P@10 |
|---|---|---|
| ONNX fp32 | 1.0000 / 1.0000 | 0.870 |
| ONNX fp16 | 1.0000 / 1.0000 | 0.870 |
| ONNX q4f16 | 0.9911 / 0.8614 | 0.865 |
| ONNX int8 | 0.9217 / 0.3807 | **0.741** |

**These figures are 300 prompts, below the 1,000-item floor §4 requires. They are historical motivation for why precision matters, not a membership determination for any pair of implementations; no `spec_id` claim in this document rests on them.**

The runtime alone is not the variable — a faithful export reproduces the space **exactly**. What breaks identity is *precision*, and it breaks it hard: 8-bit weights cost 13 points of P@10 while 4-bit weights with fp16 activations cost half a point. (An earlier revision of this section reported a flat "0.97 parity" for ONNX and drew the general conclusion that runtime substitution is unsafe. That number came from a quantized export; the generalization was wrong and is corrected here.)

"We used SigLIP" is still not a sufficient statement of identity — but the field that matters is the one that was missing from the claim, not the word ONNX.

---

## 3. The `embed_spec`

An `embed_spec` describes the deterministic function `input → vector`. It has two layers, and the split is the whole of the identity model: an `identity_basis` that is hashed into the `spec_id` (§3.1), and `implementation_records` that are recorded but **excluded from the `spec_id` hash** (§3.3). What separates them is not how much a field perturbs the output — it is whether the field is the *definition* of the space or an *account of what ran*.

The container itself is **CLOSED** per `0002` §3.1.4: exactly these three keys, all REQUIRED.

```
embed_spec = {                                    ; CLOSED: exactly these 3 keys
  identity_basis:        { ... },                 ; §3.1 — hashed to the spec_id
  realization_manifest:  { ... },                 ; §3.2 — the anchor this basis commits to
  implementation_records: [ implementation_record, ... ]   ; §3.3 — always an array
}
```

Two shape rules that exist because their absence would leave one logical spec with several byte encodings:

- **`realization_manifest` is carried here, in full.** §4 requires a membership claim to obtain and validate the anchor, and this is where it is obtained: the manifest is persisted with the `embed_spec` in the registry (§3.4), not left implicit. What remains **out of band is the probe and reference *bytes*** — the manifest holds only their digests, which are opaque to the protocol (§3.2). Without this the specification required something no wire form supplied.
- **`implementation_records` is always an array**, including when it holds one record. A singleton shorthand is NOT permitted: one record encoded two ways is two sets of Manifest bytes and two content addresses for the same logical spec, and a reader cannot tell which was intended.
- **The array MUST NOT be empty.** The realization the anchor commits to is itself an admitted implementation — it is the one that produced the committed reference outputs — so §3.3's "one record per admitted implementation" already covers it, and an `embed_spec` with no records would be one whose own anchor has no provenance. Which entry is the anchor's own is **not distinguished structurally**: nothing ties a record's `impl` back to the manifest, by design, since §3.2 keeps provenance out of the anchor. The requirement a reader can enforce is therefore non-emptiness, and this section claims no more than that.
- **The array is a set, canonically ordered.** Entries MUST be sorted ascending by the **canonical CBOR encoding of each record**, compared as byte strings, and two entries MUST NOT be byte-identical. Records are an unordered collection of admitted implementations, not an admission history: without a fixed order `[A, B]` and `[B, A]` would be two encodings of one collection, which is the same defect the singleton shorthand was refused for. No ordering carries meaning — neither admission time nor preference — and none may be inferred from position.

### 3.1 `identity_basis` (enters the `spec_id` hash)

The `identity_basis` is the hashed layer. It is the existing hard fields plus one addition: a commitment to the **reference realization** that anchors the space.

| Field | Notes |
|---|---|
| `checkpoint` | Immutable revision identifier (e.g. a commit SHA), **not** a mutable name. Hosted checkpoints are updated in place. |
| `dim` | Output dimensionality. |
| `output` | Which head/pooling is taken (e.g. `pooler_output` vs a hidden state). Different heads are different spaces. |
| `normalize` | Whether the vector is L2-normalized, and when. |
| `text_template` | The prompt wrapper for text encoding, e.g. `"This is a photo of a {}."`. A bare token and a templated one rank very differently; this is part of the function. |
| `tokenizer` | Identity, `max_length`, and padding mode. |
| `realization_commitment` | 33-byte multihash of the canonical `realization_manifest` (§3.2). The reference anchor of this identity. |

The `identity_basis` is **CLOSED** per `0002` §3.1.4: exactly these seven keys, all REQUIRED, no OPTIONAL keys, and an unknown or duplicate key MUST be rejected. Types are fixed: `dim` is a uint; `checkpoint`, `output`, `normalize` and `text_template` are tstr; `realization_commitment` is a 33-byte multihash; `tokenizer` is itself CLOSED with exactly `identity` (tstr), `max_length` (uint) and `padding` (tstr). A map that is merely *similar* to a basis is not one.

`realization_commitment` is what makes a failed-parity change **computable**. Without it, a realization that fails parity leaves every hashed field untouched, so the canonical CBOR and therefore the `spec_id` cannot change — the requirement to "mint a new `spec_id`" would have no input to act on.

### 3.2 The reference anchor: `realization_manifest`

The `realization_manifest` describes the realization whose outputs define this identity. It and **every map nested inside it** are **CLOSED** per `0002` §3.1.4: each has an exact key set, every listed key is REQUIRED, there are no OPTIONAL keys, and a reader encountering an unknown or duplicate key MUST reject the map with an explicit error. Key order is the canonical map ordering of `0002` §3.

```
realization_manifest = {                          ; CLOSED: exactly these 3 keys
  basis_core: { ... },                            ; CLOSED: exactly the 6 hard fields of §3.1
  probe_set: { count: uint, encoding: tstr, digest: bstr(33) },   ; CLOSED: exactly these 3
  reference: { count: uint, dim: uint, dtype: tstr, digest: bstr(33) }   ; CLOSED: exactly these 4
}

realization_commitment = multihash(encode_canonical(realization_manifest))
```

The CLOSED declaration covers the nested maps explicitly, as `0002` §3.1.4 requires: closing only the outer map would leave every inner map open, and the digests live in the inner maps. `basis_core`'s own nested `tokenizer` is CLOSED on the same terms as in §3.1.

**The manifest carries no implementation provenance, deliberately.** An earlier revision of this document put `runtime`, `precision`, `image_resize`, `decoder` and artifact digests inside the manifest. Because the manifest's hash *is* the `realization_commitment`, and the commitment sits inside the hashed `identity_basis`, every one of those values entered the `spec_id`: respelling `runtime` from `onnxruntime` to any synonym minted a new identity while the probe set and the reference matrix stayed byte-identical. That contradicted this document twice over — §3.3 places the implementation description outside the `spec_id` hash, and the whole point of the anchor is that identity moves on a *computed* input rather than on text a writer chose.

What defines the space is now exactly what pins the function's behaviour: the hard fields, the committed probe set, and the reference outputs those probes produced. Artifact digests are **provenance, not identity** — once `reference` pins the outputs for the committed inputs, a second realization that reproduces them belongs to the space whatever it was built from, which is precisely what §4 measures. Provenance is recorded in `implementation_record` (§3.3), which is excluded from the `spec_id` hash.

Normative constraints:

- **`basis_core` MUST be exactly the six hard fields of §3.1**, with the same values as the `identity_basis` that names this manifest — the same key set, not merely an agreeing subset. A manifest whose `basis_core` disagrees is rejected. This is what stops one basis from pointing at another basis's anchor.
- **Probe ordering and encoding.** The probe set is an **ordered sequence**. Its canonical encoding is a CBOR array of byte strings, one per probe input, in that order; `probe_set.digest` is the multihash of the `encode_canonical` of that array. `probe_set.count` MUST be ≥ 1000.
- **`probe_set.encoding` is a versioned literal**, not free text. This revision defines exactly one legal value, `"dreamdb.probe-array.v1"`, denoting the encoding above; any other value MUST be rejected. It is a hashed field, so leaving it free-form would let one logical anchor take any number of commitments.
- **Output matrix.** `dtype` is `f32le` or `f64le` — byte order is carried in the dtype, not a separate field. The matrix is row-major, and its **row order MUST match the probe-set order**. `reference.count` MUST equal `probe_set.count`; `reference.dim` MUST equal `basis_core.dim`, and both MUST be present as uints — two absent values are not a match. `reference.digest` is the multihash of the matrix bytes.

**These digests are commitments over bytes, not DreamDB Object addresses.** A conforming implementation MUST NOT resolve any of them as an address, and a garbage collector MUST NOT traverse them — to the protocol they are opaque byte strings. They therefore introduce no reference-bearing location under `0002` §3.1.0 condition 3. Whoever performs a parity certification obtains the probe and reference bytes **out of band** and MUST verify each digest before certifying. **If the bytes cannot be obtained, the certification fails closed**: the candidate does not join the identity. "The anchor is unavailable" is never grounds to admit it.

Persisting probe sets and reference outputs as DreamDB Objects is deliberately out of scope here; it would be a new reference-bearing location and needs its own format revision under `0002` §3.1.0.

**What the commitment does and does not secure.** Because `realization_commitment` is `H(realization_manifest)` and the manifest sits inside the hashed `identity_basis`, the anchor cannot be altered under a fixed `spec_id`: any edit to the declared manifest — a different probe set, a different reference matrix, a changed hard field — computes a different commitment and therefore a different identity. That is the whole of the guarantee. It does **not** establish that the committed probe or reference bytes exist or are retrievable, that any particular implementation produced them, or that a claimed measurement ever happened; and it does not stop anyone from copying an existing `identity_basis` verbatim or from misreporting a membership. Those are falsifiable by re-running the measurement against the committed bytes, not prevented by the encoding.

### 3.3 `implementation_record` (excluded from the `spec_id` hash)

One record per admitted implementation, including the anchor's own realization. A record is **CLOSED**: exactly `impl` and `parity`, both REQUIRED, with no other keys.

```
implementation_record = {                         ; CLOSED: exactly these 2 keys
  impl: {                                         ; CLOSED: exactly these 5 keys
    runtime:      tstr,
    precision:    tstr,
    image_resize: tstr,
    decoder:      tstr,
    artifacts:    [ { role: tstr, digest: bstr(33) }, ... ]   ; CLOSED entries, ordered
  },
  parity: {                                       ; CLOSED: exactly these 4 keys
    mean_f64le: bstr(8),                          ; IEEE-754 binary64 bit pattern
    min_f64le:  bstr(8),
    count:      uint,
    against:    bstr(33)
  }
}
```

`impl` binds the record to a specific build. The four descriptive fields are free text and cannot distinguish two builds on their own; the `artifacts` digests can, so a record written for one set of weights stops reading as valid once they are replaced. This makes "which implementation was measured" **falsifiable**. It is not a proof that the measurement was performed — nothing here signs or attests it — but a false record now has to name the artifacts it is false about.

**Artifact ordering.** `artifacts` MUST be sorted ascending by `role`, and within one `role` ascending by the **33 wire bytes of the digest multihash** — not by its base32 spelling, which `0002` §7.2.3 already notes is a different ordering and would put some digest pairs in the opposite sequence. Two entries MUST NOT share both `role` and `digest`. Several entries MAY share a `role`: sharded weights are one artifact set, not several roles. The list MUST include at least one entry with `role = "weights"`. Ordering is required so that two records describing the same build have **one canonical form** and compare equal. It is not a `spec_id` concern: the list is excluded from that hash, though like every other byte of a record it still affects the content address of the Manifest carrying it.

There is deliberately **no field in which a record states its own verdict**. The threshold decision is recomputed from `mean_f64le`, `min_f64le` and `count` by whoever reads the record; a `passed` flag would be a second, contradictable source for something already determined by the numbers beside it.

The `embed_spec`'s `implementation_records` array holds these, as a canonically ordered set (§3). Records never affect the `spec_id` — which is the property the anchor's provenance moving here (§3.2) restores. They do affect the content address of the Manifest that carries them, which is why their encoding is constrained at all.

**The statistics are byte patterns, not CBOR floats.** `mean_f64le` and `min_f64le` each carry the **8 bytes of the IEEE-754 binary64 value in little-endian order** — the same byte order the `f64le` matrix dtype names — inside a byte string. A reader decodes the pattern and applies the finite check and the thresholds of §4 to the recovered binary64 value.

This is not a stylistic choice. `0002` §3.1.5 forbids floating-point numbers in any field that affects hashing, and the constraint is per CBOR Object, not per identity: §3.4 records the `embed_spec` in the Manifest registry, so every byte of an `implementation_record` affects the **Manifest's** content address even though none of it affects the `spec_id`. An earlier revision of this document asserted that "no float ever enters a hashed structure" because the record is excluded from the `spec_id` — which confused *not in the `spec_id` hash* with *not in any hash*, and left the record in violation of §3.1.5.

Carrying the exact bit pattern also preserves §4's result rather than introducing a second quantization boundary: a fixed-point encoding would round the very value the threshold comparison is deciding on. A pattern encoding a NaN or an infinity MUST be rejected, as §4 already requires of the values themselves.

### 3.4 Serialization

The `embed_spec` is the CLOSED CBOR map of §3, recorded in the Manifest registry alongside the modality entry — `identity_basis`, `realization_manifest` and `implementation_records` together. The `spec_id` is the full 33-byte BLAKE3-256 multihash of the canonical CBOR encoding of the `identity_basis` only, sorted by key. Its string form is the canonical 53-character lowercase, unpadded base32 multihash from `0002` §8.1. The algorithm tag is part of that encoding; a shortened digest, naked digest, or hexadecimal prefix is not a `spec_id`.

An embedding modality MUST carry exactly one `spec=<spec_id>` parameter. Its decoded multihash MUST equal the hash of the registered `identity_basis`. Missing, repeated, malformed, truncated, or mismatching values are protocol errors.

The registered `realization_commitment` MUST equal the hash of the registered `realization_manifest` (§3.2). A registry entry failing that check is malformed, whatever its `spec_id` arithmetic says: the anchor would name a realization the entry does not carry.

The reference implementation persists the same declaration in two places that
serve different consumers, and requires them to agree:

- The embedding modality's registry value carries `embed_spec` beside the
  index binding. The modality key carries `spec=<spec_id>`. Both must be
  present together and the computed identity must equal the parameter.
- The persisted `dreamdb.schema` value uses Schema wire version 2 and carries
  `embedding_specs: { <field>: <embed_spec>, ... }`. Version 1 remains the
  exact legacy shape and is emitted when no field is identified. A v2 entry
  must name an embedding field of the same dimension; the Schema declaration,
  modality parameter and registry declaration must agree.

Schema version 2 is also the dataset-level critical-extension boundary required
by `0002` §3.1.0: readers that predate embedding identity reject the Schema
before resolving Items rather than treating the new declaration as an
ignorable annotation. A GraphIndex additionally uses its Vamana params version
2 boundary (`0013` §3.1), because a graph object can be decoded independently
of a Dataset Schema. Built-in SpatialIndexes use the same params-version split
(`0004` §3.1). Thus independently decoded index objects retain a fail-closed
boundary instead of relying only on the enclosing Schema.

No implementation may infer one copy from the other when opening malformed
data. The duplication is an integrity boundary: Schema selects the public field
while the registry binds the stored Track and indexes.

## 4. Membership: parity against the fixed anchor

A `spec_id` names a **compatibility space whose representative is its reference anchor**. Membership is decided against that anchor and nothing else:

> An implementation belongs to `spec_id` S **if and only if** it reproduces the outputs committed by S's `realization_commitment`. Encode the committed probe set — at least 1,000 items — under the candidate implementation. Against the committed reference matrix, the mean cosine MUST be ≥ **0.999** and the minimum ≥ **0.99**.

**The measurement is fully specified, because a threshold decides identity.** Two conforming implementations that disagree near the boundary would disagree about whether a `spec_id` may be reused, so the inputs and the arithmetic are pinned:

- **Inputs.** The candidate encodes exactly the committed probe set, in its committed order. The candidate matrix MUST have the same `dtype`, row count and `dim` as `reference`, and its row order MUST match the probe order. `parity.count` MUST equal `probe_set.count` and `reference.count`; a measurement over a different number of probes is not a measurement against this anchor.
- **Arithmetic.** Every value is widened to **binary64** before any arithmetic, whatever the matrix dtype. For each probe, the dot product and both squared norms are accumulated **sequentially in ascending dimension index**; `cos_i = dot / (sqrt(norm_ref) * sqrt(norm_cand))` in binary64.
- **Arithmetic discipline.** Ascending order alone does not pin the result. `0004` §5.4.1 already states this discipline for the f32 spatial path; it applies here in binary64. Every operation — each multiply, each add, the two square roots and the division — rounds **round-to-nearest, ties-to-even** in binary64. Implementations MUST NOT fuse a multiply and an add into an FMA (one rounding where the specification says two), MUST NOT re-associate or vectorize the accumulations, and MUST NOT hold intermediates at extended precision (x87 80-bit, or a wider accumulator). `-ffast-math`, `--unsafe-fp-math`, contraction defaults and equivalents MUST be disabled on this path. The canonical form is the explicit scalar loop.
- **Aggregation.** `min` is the minimum of `cos_i`. `mean` is the sum of `cos_i` accumulated **sequentially in ascending probe index**, in binary64, divided by `count`. Comparison against the thresholds is on these binary64 values, recovered from the `mean_f64le` / `min_f64le` bit patterns of §3.3 without re-rounding.
- **Degenerate values.** A NaN or infinity anywhere in either matrix, or a zero-norm row in either matrix, MUST cause the certification to be refused. Cosine is undefined for a zero vector; substituting a value would decide identity by convention.

Accumulation order is specified because floating-point addition is not associative: a reordered or vectorized sum is a different number, and near the threshold a different verdict. The same reasoning forces the rest of the discipline — an opportunistic FMA rounds once where two roundings were specified, and an extended-precision accumulator carries bits the binary64 result should have dropped.

**Certification requires the anchor's manifest.** A `spec_id` may be asserted on its own — the identity arithmetic needs only the `identity_basis`. But **a membership claim MUST NOT be made without validating the anchor's `realization_manifest`**, because every term of the measurement is defined by it: the probe set, its committed count, the reference matrix, its dtype and dim. The manifest travels in the `embed_spec` itself (§3), so it is always available where the records are; what must still be fetched out of band are the probe and reference **bytes** its digests name. A certification that cannot obtain those bytes **fails closed**; it does not fall back to checking the threshold alone. In particular `parity.count` is meaningless without the committed count it must equal, and two records certified against one anchor necessarily state the same `count`.

This is **set membership against a fixed representative**, not a tolerance relation between implementations. The distinction is the point:

- **Chained certification is forbidden.** An implementation MUST NOT acquire a `spec_id` by passing parity against another implementation that already holds it. The only permitted comparison is against the anchor.
- Were chaining allowed, `A ≈ B` and `B ≈ C` with `A ≉ C` would put mutually incomparable implementations in one identity. Under anchor-relative membership that state is not representable: B and C are each measured against A's committed outputs, and each is either in or out.

The rule needs no machinery of its own. `parity.against` MUST equal the basis's `realization_commitment` byte for byte; anything else is refused. A chained certification names some other implementation's anchor and is refused by that same comparison — there is no separate notion of an implementation-record identity in the protocol, and inventing one to give chaining its own error would add a concept the format does not have.

### 4.1 Failing parity mints a new anchor

When implementation X fails the measurement against S's anchor, X does not join S. X MAY instead become the reference anchor of a **new** `identity_basis` S′:

1. S′ MAY carry hard fields **identical** to S's, field for field.
2. S′'s `realization_commitment` **necessarily differs** — its `realization_manifest` commits to X's own reference outputs. (It carries no implementation descriptor; §3.2 excludes provenance from the manifest deliberately.)
3. Therefore S′'s canonical CBOR differs, and its `spec_id` differs **deterministically**.
4. A later implementation joins S′ only by passing parity against **S′'s** anchor.

This closes the gap the identity model previously had: there is now a specification-computable input that a failed realization change moves.

**A consequence that must not be discovered by surprise.** The first realization of a hard-field map is necessarily its own anchor, so **identical hard fields no longer imply an identical `spec_id`**. Two independently built implementations that never measured against each other receive two different identities. This is intended — sameness is now measured rather than asserted. The route to a shared identity is explicit: to use an existing `spec_id`, measure against its anchor; an implementation that does not is minting its own.

### 4.2 What membership does and does not establish

Passing the measurement establishes membership **as this specification defines it**. It is not a claim that two implementations compute the same function over the whole input domain, and it does not establish that the probe set resists overfitting by an implementation tuned to it. A fixed probe set can only demonstrate agreement on the probes it contains.

Accordingly, the correct statement about a substituted checkpoint is:

> If an implementation — including one built on a different checkpoint — passes the stated measurement against the fixed anchor, it MAY join that identity class.

Not that it *is* the same function.

### 4.3 Retrieval evidence does not alter identity

A change that fails parity is a different identity. Retrieval measurements MUST NOT be used to retain the original `spec_id` for an implementation that failed parity — an earlier revision of this section permitted exactly that, and it contradicted §7, which holds comparability and retrieval quality to be independent.

Retrieval evidence keeps its uses, both outside identity: an operator weighs it when deciding whether to declare `compatible_with` (`0017` §2.2) between two distinct `spec_id`s, and when deciding whether a re-encode is worth its cost. Note that `compatible_with` carries `modality`, `relationship`, `coverage` and `transform_ref` — **it has no field in which a measurement is recorded**. The evidence informs the operator's decision; it is not written into the structure, and it never rewrites identity equality.

Implementations MUST NOT assume export equivalence without measuring it, and MUST measure the precision they actually ship rather than the reference export.

## 5. Binding: the `spec_id` lives in the modality string

The complete base32 multihash is carried as a modality parameter. This example is also grammatical under `0002` §5.1: human-readable parameter values use underscores rather than forbidden hyphens.

```
embedding.f32.dim=256.model=siglip2_b16_256.spec=d2s7t6cwkmc2ns26nrg4au3kq5u44pqdvxljwumdoa2xhrvnoffju.bucketed
```

The exact `identity_basis` that produces this example ID — including its `realization_commitment` — is frozen in the `0024.embedding-spec.identity.001` conformance vector; the string is not an illustrative truncation.

`model` is a human-readable convenience and carries no semantics. `spec` is authoritative.

This binding is deliberate and has three consequences that fall out for free:

1. **Different specs are different modalities**, therefore different Tracks, therefore different SpatialIndexes. Incompatible vectors cannot occupy one field. Multi-table probing (`spatial_index_hashes`) and cross-dataset joint search naturally decline to match.
2. The existing `algo_eq` / `key_eq` comparison machinery applies unchanged.
3. Coexistence during migration is expressed by `compatible_with` (spec/0017 §2.2), which already carries a `coverage` of `partial` or `complete`.

**Rationale for binding to an immutable modality rather than a mutable collection config.** Systems that attach the encoder to a mutable collection setting share a documented hazard: changing the setting does not re-encode existing rows, so old and new vectors coexist in one field with no marker distinguishing them. Binding to the modality makes that state unrepresentable — the old vectors remain on the old Track, addressed by the old modality, and a reader must name which one it wants.

---

## 6. Validation is mandatory and MUST fail loudly

A recorded invariant that nothing checks is not an invariant. Implementations MUST validate at these points:

**6.1 Query.** A caller supplying a query vector MUST supply the `spec_id` it was produced under, and it MUST equal the `spec_id` of the target modality. If it differs the query MUST be refused. The error MUST name both `spec_id`s. When the caller also supplies the complete conflicting `embed_spec`, the error MUST additionally name the first `identity_basis` field that differs; an API carrying only the content ID cannot reconstruct that field and MUST NOT guess it.

**A `compatible_with` declaration does not relax this.** It never authorizes scoring a vector produced under one `spec_id` against an index built under another — that is precisely the silent, plausible, total failure §2 describes, and §4.3 has already established that `compatible_with` cannot rewrite identity equality. What it authorizes is upstream of the vector: a logical-query planner holding the caller's **raw input** MAY encode that input separately under each declared version and fuse the result sets, exactly as `0017` §4.2 describes. Each encoding is then scored against the index of its own `spec_id`, and no vector ever crosses a spec boundary.

**6.2 Append.** Writing vectors into a modality whose `spec_id` differs from the writer's MUST be refused.

Public append and query APIs distinguish identified and legacy calls. An
identified field requires the caller's exact ID; a legacy field rejects an
unexpected ID. The legacy overloads remain valid only for fields whose Schema
and registry both omit identity. This preserves existing datasets without ever
inventing an identity for their stored vectors.

**6.3 Multi-table and joint search.** All participating tables MUST share one `spec_id`. Mixed-spec probing MUST be refused, not silently ranked.

**6.4 Index build.** Training centroids or a graph MUST record the `spec_id` of the vectors trained on, and MUST refuse to publish a SpatialIndex or GraphIndex whose `spec_id` differs from the modality it is bound to. `0004` §3.1 and `0013` §3.1 encode this as the optional `embedding_spec_id` field: it is required for an identified modality and omitted only for a legacy one. Every table in a multi-table binding and every descendant produced by index maintenance carries the same value.

In every case the required behaviour is refusal. Degrading to a best-effort answer reproduces exactly the failure this document exists to prevent.

---

## 7. Relationship to acceptance gating

`spec_id` equality proves two vector sets are *comparable*. It proves nothing about whether retrieval is *good*. The two checks are independent and both are required before promoting an index. §4.3 states the same separation from the identity side: retrieval evidence never buys back an identity that parity refused.

- **Comparability** — this document, §6.
- **Retrieval** — self-retrieval and ground-truth relevance gates, run against the built index before it is published.

An index may pass §6 perfectly and still retrieve badly (a well-formed index over a poorly-chosen model). It may also retrieve plausibly while violating §6 (the silent-mismatch case). Neither check substitutes for the other.

---

## 8. Migration and the superseded draft spelling

Unchanged from spec/0017. A new `spec_id` — whether from changed hard fields or from a new reference anchor under §4.1 — is a new modality; `compatible_with` declares the supersedes relationship and its `coverage`; `reencode_state` checkpoints the bulk transform. This document adds only the requirement that the new modality's `spec_id` be recorded at creation, so the migration boundary is machine-checkable rather than documented.

The earlier draft's two example spellings have no implicit compatibility:

- `model=siglip2-b16-256` is not a modality at all under `0002` §5.1; the reference parser has always rejected its hyphens.
- a syntactically valid `spec=a3f9c1` is only a 24-bit prefix. It cannot identify or reconstruct a unique `identity_basis` and MUST NOT compare equal to any full `spec_id`.

At the time this correction was adopted, the reference implementation did not emit or enforce either draft spelling. If an external client nevertheless persisted a modality containing a six-hex `spec` value, an `0024` reader treats it as a legacy unidentified embedding, not as an alias. Migration requires registering the complete `embed_spec`, minting the new full-ID modality, and re-encoding through `0017`; writers and readers MUST NOT rewrite an existing modality or guess the missing hash bits. Modalities without a `spec` parameter likewise remain pre-`0024` identities and receive no implicit mapping.
