DreamDB

Spec 0024 — Embedding Spec Identity

Status: Implemented. Depends on: spec/0001, spec/0002, spec/0004, spec/0015, spec/0017. Motivation: Two vectors of the same dimension are silently comparable. Nothing in the protocol prevents a query encoded by one function from being scored against an index built by another — cosine similarity computes fine, scores look plausible, and the results are confidently wrong. This has already happened in practice, and it is not detectable from the data. spec/0017 makes migration between embedding versions graceful; this document defines what makes two embeddings the same version in the first place, and requires that mismatches fail loudly.


1. Purpose

spec/0017 defines versioned modality strings and the compatible_with hint so old and new tracks can coexist through a migration. It explicitly does not define what makes a vector belong to one version rather than another — it treats that as operator policy.

That gap is a correctness hole, not a policy question. This document closes it:

  • The embed_spec: the complete, pinned description of the deterministic function that produced a vector.
  • The spec_id: a content hash over the parts of that description which, if changed, produce an incomparable space.
  • Modality binding: the spec_id travels in the modality string, so incompatible embeddings cannot occupy the same field.
  • Mandatory validation: where the check happens and what it must refuse.

What this document does NOT define:

  • Which model to use. Operator policy, as in spec/0017.
  • Whether a model is good. Retrieval quality is an evaluation concern; see the acceptance-gate discipline in §6.
  • Storage layout. Unchanged — spec/0007, spec/0010.

2. The problem is that dimension is not identity

A 768-d SigLIP vector and a 768-d CLIP ViT-L/14 vector are structurally indistinguishable. Every layer of the system accepts both:

  • Dataset::query_vector_scored validates query.len() == schema.dim and nothing more.
  • IvfCosine::hash_vector computes dot products against centroids regardless of provenance.
  • The rerank pass computes exact cosine on f32 that means something different.

The failure is silent, plausible, and total. There is no error, the score distribution looks normal, and every returned record is wrong.

Identity is the function, not the model. Measured evidence, google/siglip-base-patch16-224 text tower, 300 prompts, pooler-output cosine against the PyTorch source:

exportparity (mean / min)label-as-query P@10
ONNX fp321.0000 / 1.00000.870
ONNX fp161.0000 / 1.00000.870
ONNX q4f160.9911 / 0.86140.865
ONNX int80.9217 / 0.38070.741

These figures are 300 prompts, below the 1,000-item floor §4 requires. They are historical motivation for why precision matters, not a membership determination for any pair of implementations; no spec_id claim in this document rests on them.

The runtime alone is not the variable — a faithful export reproduces the space exactly. What breaks identity is precision, and it breaks it hard: 8-bit weights cost 13 points of P@10 while 4-bit weights with fp16 activations cost half a point. (An earlier revision of this section reported a flat "0.97 parity" for ONNX and drew the general conclusion that runtime substitution is unsafe. That number came from a quantized export; the generalization was wrong and is corrected here.)

"We used SigLIP" is still not a sufficient statement of identity — but the field that matters is the one that was missing from the claim, not the word ONNX.


3. The embed_spec

An embed_spec describes the deterministic function input → vector. It has two layers, and the split is the whole of the identity model: an identity_basis that is hashed into the spec_id (§3.1), and implementation_records that are recorded but excluded from the spec_id hash (§3.3). What separates them is not how much a field perturbs the output — it is whether the field is the definition of the space or an account of what ran.

The container itself is CLOSED per 0002 §3.1.4: exactly these three keys, all REQUIRED.

embed_spec = {                                    ; CLOSED: exactly these 3 keys
  identity_basis:        { ... },                 ; §3.1 — hashed to the spec_id
  realization_manifest:  { ... },                 ; §3.2 — the anchor this basis commits to
  implementation_records: [ implementation_record, ... ]   ; §3.3 — always an array
}

Two shape rules that exist because their absence would leave one logical spec with several byte encodings:

  • realization_manifest is carried here, in full. §4 requires a membership claim to obtain and validate the anchor, and this is where it is obtained: the manifest is persisted with the embed_spec in the registry (§3.4), not left implicit. What remains out of band is the probe and reference bytes — the manifest holds only their digests, which are opaque to the protocol (§3.2). Without this the specification required something no wire form supplied.
  • implementation_records is always an array, including when it holds one record. A singleton shorthand is NOT permitted: one record encoded two ways is two sets of Manifest bytes and two content addresses for the same logical spec, and a reader cannot tell which was intended.
  • The array MUST NOT be empty. The realization the anchor commits to is itself an admitted implementation — it is the one that produced the committed reference outputs — so §3.3's "one record per admitted implementation" already covers it, and an embed_spec with no records would be one whose own anchor has no provenance. Which entry is the anchor's own is not distinguished structurally: nothing ties a record's impl back to the manifest, by design, since §3.2 keeps provenance out of the anchor. The requirement a reader can enforce is therefore non-emptiness, and this section claims no more than that.
  • The array is a set, canonically ordered. Entries MUST be sorted ascending by the canonical CBOR encoding of each record, compared as byte strings, and two entries MUST NOT be byte-identical. Records are an unordered collection of admitted implementations, not an admission history: without a fixed order [A, B] and [B, A] would be two encodings of one collection, which is the same defect the singleton shorthand was refused for. No ordering carries meaning — neither admission time nor preference — and none may be inferred from position.

3.1 identity_basis (enters the spec_id hash)

The identity_basis is the hashed layer. It is the existing hard fields plus one addition: a commitment to the reference realization that anchors the space.

FieldNotes
checkpointImmutable revision identifier (e.g. a commit SHA), not a mutable name. Hosted checkpoints are updated in place.
dimOutput dimensionality.
outputWhich head/pooling is taken (e.g. pooler_output vs a hidden state). Different heads are different spaces.
normalizeWhether the vector is L2-normalized, and when.
text_templateThe prompt wrapper for text encoding, e.g. "This is a photo of a {}.". A bare token and a templated one rank very differently; this is part of the function.
tokenizerIdentity, max_length, and padding mode.
realization_commitment33-byte multihash of the canonical realization_manifest (§3.2). The reference anchor of this identity.

The identity_basis is CLOSED per 0002 §3.1.4: exactly these seven keys, all REQUIRED, no OPTIONAL keys, and an unknown or duplicate key MUST be rejected. Types are fixed: dim is a uint; checkpoint, output, normalize and text_template are tstr; realization_commitment is a 33-byte multihash; tokenizer is itself CLOSED with exactly identity (tstr), max_length (uint) and padding (tstr). A map that is merely similar to a basis is not one.

realization_commitment is what makes a failed-parity change computable. Without it, a realization that fails parity leaves every hashed field untouched, so the canonical CBOR and therefore the spec_id cannot change — the requirement to "mint a new spec_id" would have no input to act on.

3.2 The reference anchor: realization_manifest

The realization_manifest describes the realization whose outputs define this identity. It and every map nested inside it are CLOSED per 0002 §3.1.4: each has an exact key set, every listed key is REQUIRED, there are no OPTIONAL keys, and a reader encountering an unknown or duplicate key MUST reject the map with an explicit error. Key order is the canonical map ordering of 0002 §3.

realization_manifest = {                          ; CLOSED: exactly these 3 keys
  basis_core: { ... },                            ; CLOSED: exactly the 6 hard fields of §3.1
  probe_set: { count: uint, encoding: tstr, digest: bstr(33) },   ; CLOSED: exactly these 3
  reference: { count: uint, dim: uint, dtype: tstr, digest: bstr(33) }   ; CLOSED: exactly these 4
}

realization_commitment = multihash(encode_canonical(realization_manifest))

The CLOSED declaration covers the nested maps explicitly, as 0002 §3.1.4 requires: closing only the outer map would leave every inner map open, and the digests live in the inner maps. basis_core's own nested tokenizer is CLOSED on the same terms as in §3.1.

The manifest carries no implementation provenance, deliberately. An earlier revision of this document put runtime, precision, image_resize, decoder and artifact digests inside the manifest. Because the manifest's hash is the realization_commitment, and the commitment sits inside the hashed identity_basis, every one of those values entered the spec_id: respelling runtime from onnxruntime to any synonym minted a new identity while the probe set and the reference matrix stayed byte-identical. That contradicted this document twice over — §3.3 places the implementation description outside the spec_id hash, and the whole point of the anchor is that identity moves on a computed input rather than on text a writer chose.

What defines the space is now exactly what pins the function's behaviour: the hard fields, the committed probe set, and the reference outputs those probes produced. Artifact digests are provenance, not identity — once reference pins the outputs for the committed inputs, a second realization that reproduces them belongs to the space whatever it was built from, which is precisely what §4 measures. Provenance is recorded in implementation_record (§3.3), which is excluded from the spec_id hash.

Normative constraints:

  • basis_core MUST be exactly the six hard fields of §3.1, with the same values as the identity_basis that names this manifest — the same key set, not merely an agreeing subset. A manifest whose basis_core disagrees is rejected. This is what stops one basis from pointing at another basis's anchor.
  • Probe ordering and encoding. The probe set is an ordered sequence. Its canonical encoding is a CBOR array of byte strings, one per probe input, in that order; probe_set.digest is the multihash of the encode_canonical of that array. probe_set.count MUST be ≥ 1000.
  • probe_set.encoding is a versioned literal, not free text. This revision defines exactly one legal value, "dreamdb.probe-array.v1", denoting the encoding above; any other value MUST be rejected. It is a hashed field, so leaving it free-form would let one logical anchor take any number of commitments.
  • Output matrix. dtype is f32le or f64le — byte order is carried in the dtype, not a separate field. The matrix is row-major, and its row order MUST match the probe-set order. reference.count MUST equal probe_set.count; reference.dim MUST equal basis_core.dim, and both MUST be present as uints — two absent values are not a match. reference.digest is the multihash of the matrix bytes.

These digests are commitments over bytes, not DreamDB Object addresses. A conforming implementation MUST NOT resolve any of them as an address, and a garbage collector MUST NOT traverse them — to the protocol they are opaque byte strings. They therefore introduce no reference-bearing location under 0002 §3.1.0 condition 3. Whoever performs a parity certification obtains the probe and reference bytes out of band and MUST verify each digest before certifying. If the bytes cannot be obtained, the certification fails closed: the candidate does not join the identity. "The anchor is unavailable" is never grounds to admit it.

Persisting probe sets and reference outputs as DreamDB Objects is deliberately out of scope here; it would be a new reference-bearing location and needs its own format revision under 0002 §3.1.0.

What the commitment does and does not secure. Because realization_commitment is H(realization_manifest) and the manifest sits inside the hashed identity_basis, the anchor cannot be altered under a fixed spec_id: any edit to the declared manifest — a different probe set, a different reference matrix, a changed hard field — computes a different commitment and therefore a different identity. That is the whole of the guarantee. It does not establish that the committed probe or reference bytes exist or are retrievable, that any particular implementation produced them, or that a claimed measurement ever happened; and it does not stop anyone from copying an existing identity_basis verbatim or from misreporting a membership. Those are falsifiable by re-running the measurement against the committed bytes, not prevented by the encoding.

3.3 implementation_record (excluded from the spec_id hash)

One record per admitted implementation, including the anchor's own realization. A record is CLOSED: exactly impl and parity, both REQUIRED, with no other keys.

implementation_record = {                         ; CLOSED: exactly these 2 keys
  impl: {                                         ; CLOSED: exactly these 5 keys
    runtime:      tstr,
    precision:    tstr,
    image_resize: tstr,
    decoder:      tstr,
    artifacts:    [ { role: tstr, digest: bstr(33) }, ... ]   ; CLOSED entries, ordered
  },
  parity: {                                       ; CLOSED: exactly these 4 keys
    mean_f64le: bstr(8),                          ; IEEE-754 binary64 bit pattern
    min_f64le:  bstr(8),
    count:      uint,
    against:    bstr(33)
  }
}

impl binds the record to a specific build. The four descriptive fields are free text and cannot distinguish two builds on their own; the artifacts digests can, so a record written for one set of weights stops reading as valid once they are replaced. This makes "which implementation was measured" falsifiable. It is not a proof that the measurement was performed — nothing here signs or attests it — but a false record now has to name the artifacts it is false about.

Artifact ordering. artifacts MUST be sorted ascending by role, and within one role ascending by the 33 wire bytes of the digest multihash — not by its base32 spelling, which 0002 §7.2.3 already notes is a different ordering and would put some digest pairs in the opposite sequence. Two entries MUST NOT share both role and digest. Several entries MAY share a role: sharded weights are one artifact set, not several roles. The list MUST include at least one entry with role = "weights". Ordering is required so that two records describing the same build have one canonical form and compare equal. It is not a spec_id concern: the list is excluded from that hash, though like every other byte of a record it still affects the content address of the Manifest carrying it.

There is deliberately no field in which a record states its own verdict. The threshold decision is recomputed from mean_f64le, min_f64le and count by whoever reads the record; a passed flag would be a second, contradictable source for something already determined by the numbers beside it.

The embed_spec's implementation_records array holds these, as a canonically ordered set (§3). Records never affect the spec_id — which is the property the anchor's provenance moving here (§3.2) restores. They do affect the content address of the Manifest that carries them, which is why their encoding is constrained at all.

The statistics are byte patterns, not CBOR floats. mean_f64le and min_f64le each carry the 8 bytes of the IEEE-754 binary64 value in little-endian order — the same byte order the f64le matrix dtype names — inside a byte string. A reader decodes the pattern and applies the finite check and the thresholds of §4 to the recovered binary64 value.

This is not a stylistic choice. 0002 §3.1.5 forbids floating-point numbers in any field that affects hashing, and the constraint is per CBOR Object, not per identity: §3.4 records the embed_spec in the Manifest registry, so every byte of an implementation_record affects the Manifest's content address even though none of it affects the spec_id. An earlier revision of this document asserted that "no float ever enters a hashed structure" because the record is excluded from the spec_id — which confused not in the spec_id hash with not in any hash, and left the record in violation of §3.1.5.

Carrying the exact bit pattern also preserves §4's result rather than introducing a second quantization boundary: a fixed-point encoding would round the very value the threshold comparison is deciding on. A pattern encoding a NaN or an infinity MUST be rejected, as §4 already requires of the values themselves.

3.4 Serialization

The embed_spec is the CLOSED CBOR map of §3, recorded in the Manifest registry alongside the modality entry — identity_basis, realization_manifest and implementation_records together. The spec_id is the full 33-byte BLAKE3-256 multihash of the canonical CBOR encoding of the identity_basis only, sorted by key. Its string form is the canonical 53-character lowercase, unpadded base32 multihash from 0002 §8.1. The algorithm tag is part of that encoding; a shortened digest, naked digest, or hexadecimal prefix is not a spec_id.

An embedding modality MUST carry exactly one spec=<spec_id> parameter. Its decoded multihash MUST equal the hash of the registered identity_basis. Missing, repeated, malformed, truncated, or mismatching values are protocol errors.

The registered realization_commitment MUST equal the hash of the registered realization_manifest (§3.2). A registry entry failing that check is malformed, whatever its spec_id arithmetic says: the anchor would name a realization the entry does not carry.

The reference implementation persists the same declaration in two places that serve different consumers, and requires them to agree:

  • The embedding modality's registry value carries embed_spec beside the index binding. The modality key carries spec=<spec_id>. Both must be present together and the computed identity must equal the parameter.
  • The persisted dreamdb.schema value uses Schema wire version 2 and carries embedding_specs: { <field>: <embed_spec>, ... }. Version 1 remains the exact legacy shape and is emitted when no field is identified. A v2 entry must name an embedding field of the same dimension; the Schema declaration, modality parameter and registry declaration must agree.

Schema version 2 is also the dataset-level critical-extension boundary required by 0002 §3.1.0: readers that predate embedding identity reject the Schema before resolving Items rather than treating the new declaration as an ignorable annotation. A GraphIndex additionally uses its Vamana params version 2 boundary (0013 §3.1), because a graph object can be decoded independently of a Dataset Schema. Built-in SpatialIndexes use the same params-version split (0004 §3.1). Thus independently decoded index objects retain a fail-closed boundary instead of relying only on the enclosing Schema.

No implementation may infer one copy from the other when opening malformed data. The duplication is an integrity boundary: Schema selects the public field while the registry binds the stored Track and indexes.

4. Membership: parity against the fixed anchor

A spec_id names a compatibility space whose representative is its reference anchor. Membership is decided against that anchor and nothing else:

An implementation belongs to spec_id S if and only if it reproduces the outputs committed by S's realization_commitment. Encode the committed probe set — at least 1,000 items — under the candidate implementation. Against the committed reference matrix, the mean cosine MUST be ≥ 0.999 and the minimum ≥ 0.99.

The measurement is fully specified, because a threshold decides identity. Two conforming implementations that disagree near the boundary would disagree about whether a spec_id may be reused, so the inputs and the arithmetic are pinned:

  • Inputs. The candidate encodes exactly the committed probe set, in its committed order. The candidate matrix MUST have the same dtype, row count and dim as reference, and its row order MUST match the probe order. parity.count MUST equal probe_set.count and reference.count; a measurement over a different number of probes is not a measurement against this anchor.
  • Arithmetic. Every value is widened to binary64 before any arithmetic, whatever the matrix dtype. For each probe, the dot product and both squared norms are accumulated sequentially in ascending dimension index; cos_i = dot / (sqrt(norm_ref) * sqrt(norm_cand)) in binary64.
  • Arithmetic discipline. Ascending order alone does not pin the result. 0004 §5.4.1 already states this discipline for the f32 spatial path; it applies here in binary64. Every operation — each multiply, each add, the two square roots and the division — rounds round-to-nearest, ties-to-even in binary64. Implementations MUST NOT fuse a multiply and an add into an FMA (one rounding where the specification says two), MUST NOT re-associate or vectorize the accumulations, and MUST NOT hold intermediates at extended precision (x87 80-bit, or a wider accumulator). -ffast-math, --unsafe-fp-math, contraction defaults and equivalents MUST be disabled on this path. The canonical form is the explicit scalar loop.
  • Aggregation. min is the minimum of cos_i. mean is the sum of cos_i accumulated sequentially in ascending probe index, in binary64, divided by count. Comparison against the thresholds is on these binary64 values, recovered from the mean_f64le / min_f64le bit patterns of §3.3 without re-rounding.
  • Degenerate values. A NaN or infinity anywhere in either matrix, or a zero-norm row in either matrix, MUST cause the certification to be refused. Cosine is undefined for a zero vector; substituting a value would decide identity by convention.

Accumulation order is specified because floating-point addition is not associative: a reordered or vectorized sum is a different number, and near the threshold a different verdict. The same reasoning forces the rest of the discipline — an opportunistic FMA rounds once where two roundings were specified, and an extended-precision accumulator carries bits the binary64 result should have dropped.

Certification requires the anchor's manifest. A spec_id may be asserted on its own — the identity arithmetic needs only the identity_basis. But a membership claim MUST NOT be made without validating the anchor's realization_manifest, because every term of the measurement is defined by it: the probe set, its committed count, the reference matrix, its dtype and dim. The manifest travels in the embed_spec itself (§3), so it is always available where the records are; what must still be fetched out of band are the probe and reference bytes its digests name. A certification that cannot obtain those bytes fails closed; it does not fall back to checking the threshold alone. In particular parity.count is meaningless without the committed count it must equal, and two records certified against one anchor necessarily state the same count.

This is set membership against a fixed representative, not a tolerance relation between implementations. The distinction is the point:

  • Chained certification is forbidden. An implementation MUST NOT acquire a spec_id by passing parity against another implementation that already holds it. The only permitted comparison is against the anchor.
  • Were chaining allowed, A ≈ B and B ≈ C with A ≉ C would put mutually incomparable implementations in one identity. Under anchor-relative membership that state is not representable: B and C are each measured against A's committed outputs, and each is either in or out.

The rule needs no machinery of its own. parity.against MUST equal the basis's realization_commitment byte for byte; anything else is refused. A chained certification names some other implementation's anchor and is refused by that same comparison — there is no separate notion of an implementation-record identity in the protocol, and inventing one to give chaining its own error would add a concept the format does not have.

4.1 Failing parity mints a new anchor

When implementation X fails the measurement against S's anchor, X does not join S. X MAY instead become the reference anchor of a new identity_basis S′:

  1. S′ MAY carry hard fields identical to S's, field for field.
  2. S′'s realization_commitment necessarily differs — its realization_manifest commits to X's own reference outputs. (It carries no implementation descriptor; §3.2 excludes provenance from the manifest deliberately.)
  3. Therefore S′'s canonical CBOR differs, and its spec_id differs deterministically.
  4. A later implementation joins S′ only by passing parity against S′'s anchor.

This closes the gap the identity model previously had: there is now a specification-computable input that a failed realization change moves.

A consequence that must not be discovered by surprise. The first realization of a hard-field map is necessarily its own anchor, so identical hard fields no longer imply an identical spec_id. Two independently built implementations that never measured against each other receive two different identities. This is intended — sameness is now measured rather than asserted. The route to a shared identity is explicit: to use an existing spec_id, measure against its anchor; an implementation that does not is minting its own.

4.2 What membership does and does not establish

Passing the measurement establishes membership as this specification defines it. It is not a claim that two implementations compute the same function over the whole input domain, and it does not establish that the probe set resists overfitting by an implementation tuned to it. A fixed probe set can only demonstrate agreement on the probes it contains.

Accordingly, the correct statement about a substituted checkpoint is:

If an implementation — including one built on a different checkpoint — passes the stated measurement against the fixed anchor, it MAY join that identity class.

Not that it is the same function.

4.3 Retrieval evidence does not alter identity

A change that fails parity is a different identity. Retrieval measurements MUST NOT be used to retain the original spec_id for an implementation that failed parity — an earlier revision of this section permitted exactly that, and it contradicted §7, which holds comparability and retrieval quality to be independent.

Retrieval evidence keeps its uses, both outside identity: an operator weighs it when deciding whether to declare compatible_with (0017 §2.2) between two distinct spec_ids, and when deciding whether a re-encode is worth its cost. Note that compatible_with carries modality, relationship, coverage and transform_ref — it has no field in which a measurement is recorded. The evidence informs the operator's decision; it is not written into the structure, and it never rewrites identity equality.

Implementations MUST NOT assume export equivalence without measuring it, and MUST measure the precision they actually ship rather than the reference export.

5. Binding: the spec_id lives in the modality string

The complete base32 multihash is carried as a modality parameter. This example is also grammatical under 0002 §5.1: human-readable parameter values use underscores rather than forbidden hyphens.

embedding.f32.dim=256.model=siglip2_b16_256.spec=d2s7t6cwkmc2ns26nrg4au3kq5u44pqdvxljwumdoa2xhrvnoffju.bucketed

The exact identity_basis that produces this example ID — including its realization_commitment — is frozen in the 0024.embedding-spec.identity.001 conformance vector; the string is not an illustrative truncation.

model is a human-readable convenience and carries no semantics. spec is authoritative.

This binding is deliberate and has three consequences that fall out for free:

  1. Different specs are different modalities, therefore different Tracks, therefore different SpatialIndexes. Incompatible vectors cannot occupy one field. Multi-table probing (spatial_index_hashes) and cross-dataset joint search naturally decline to match.
  2. The existing algo_eq / key_eq comparison machinery applies unchanged.
  3. Coexistence during migration is expressed by compatible_with (spec/0017 §2.2), which already carries a coverage of partial or complete.

Rationale for binding to an immutable modality rather than a mutable collection config. Systems that attach the encoder to a mutable collection setting share a documented hazard: changing the setting does not re-encode existing rows, so old and new vectors coexist in one field with no marker distinguishing them. Binding to the modality makes that state unrepresentable — the old vectors remain on the old Track, addressed by the old modality, and a reader must name which one it wants.


6. Validation is mandatory and MUST fail loudly

A recorded invariant that nothing checks is not an invariant. Implementations MUST validate at these points:

6.1 Query. A caller supplying a query vector MUST supply the spec_id it was produced under, and it MUST equal the spec_id of the target modality. If it differs the query MUST be refused. The error MUST name both spec_ids. When the caller also supplies the complete conflicting embed_spec, the error MUST additionally name the first identity_basis field that differs; an API carrying only the content ID cannot reconstruct that field and MUST NOT guess it.

A compatible_with declaration does not relax this. It never authorizes scoring a vector produced under one spec_id against an index built under another — that is precisely the silent, plausible, total failure §2 describes, and §4.3 has already established that compatible_with cannot rewrite identity equality. What it authorizes is upstream of the vector: a logical-query planner holding the caller's raw input MAY encode that input separately under each declared version and fuse the result sets, exactly as 0017 §4.2 describes. Each encoding is then scored against the index of its own spec_id, and no vector ever crosses a spec boundary.

6.2 Append. Writing vectors into a modality whose spec_id differs from the writer's MUST be refused.

Public append and query APIs distinguish identified and legacy calls. An identified field requires the caller's exact ID; a legacy field rejects an unexpected ID. The legacy overloads remain valid only for fields whose Schema and registry both omit identity. This preserves existing datasets without ever inventing an identity for their stored vectors.

6.3 Multi-table and joint search. All participating tables MUST share one spec_id. Mixed-spec probing MUST be refused, not silently ranked.

6.4 Index build. Training centroids or a graph MUST record the spec_id of the vectors trained on, and MUST refuse to publish a SpatialIndex or GraphIndex whose spec_id differs from the modality it is bound to. 0004 §3.1 and 0013 §3.1 encode this as the optional embedding_spec_id field: it is required for an identified modality and omitted only for a legacy one. Every table in a multi-table binding and every descendant produced by index maintenance carries the same value.

In every case the required behaviour is refusal. Degrading to a best-effort answer reproduces exactly the failure this document exists to prevent.


7. Relationship to acceptance gating

spec_id equality proves two vector sets are comparable. It proves nothing about whether retrieval is good. The two checks are independent and both are required before promoting an index. §4.3 states the same separation from the identity side: retrieval evidence never buys back an identity that parity refused.

  • Comparability — this document, §6.
  • Retrieval — self-retrieval and ground-truth relevance gates, run against the built index before it is published.

An index may pass §6 perfectly and still retrieve badly (a well-formed index over a poorly-chosen model). It may also retrieve plausibly while violating §6 (the silent-mismatch case). Neither check substitutes for the other.


8. Migration and the superseded draft spelling

Unchanged from spec/0017. A new spec_id — whether from changed hard fields or from a new reference anchor under §4.1 — is a new modality; compatible_with declares the supersedes relationship and its coverage; reencode_state checkpoints the bulk transform. This document adds only the requirement that the new modality's spec_id be recorded at creation, so the migration boundary is machine-checkable rather than documented.

The earlier draft's two example spellings have no implicit compatibility:

  • model=siglip2-b16-256 is not a modality at all under 0002 §5.1; the reference parser has always rejected its hyphens.
  • a syntactically valid spec=a3f9c1 is only a 24-bit prefix. It cannot identify or reconstruct a unique identity_basis and MUST NOT compare equal to any full spec_id.

At the time this correction was adopted, the reference implementation did not emit or enforce either draft spelling. If an external client nevertheless persisted a modality containing a six-hex spec value, an 0024 reader treats it as a legacy unidentified embedding, not as an alias. Migration requires registering the complete embed_spec, minting the new full-ID modality, and re-encoding through 0017; writers and readers MUST NOT rewrite an existing modality or guess the missing hash bits. Modalities without a spec parameter likewise remain pre-0024 identities and receive no implicit mapping.