DreamDB

Spec 0010 — Vector Compression (Bucket-Internal Quantization)

Status: Normative (v0.1) for the VectorCompressor framework + the dreamdb.raw-f32, dreamdb.rabitq-cosine, and dreamdb.pq-cosine algorithms. Implemented in dreamdb-protocol vector_compressor (encode/decode/ADC), both inline and reference forms of the 200-byte compressed bucket header, and the rerank query path (dreamdb-dataset iter: Schema-decided exactness, one approximate scale for pool selection, then grouped exact-source fetches). Gated by conformance at 0009 §8.5.1 (dreamdb-conformance/vectors/0010/: VectorCompressor CBOR roundtrip for all three algorithms, RaBitQ encode-determinism vs the scalar-canonical code, raw-f32 identity code, and the 200-byte v2 bucket header). dreamdb.qinco-cosine is reserved but deferred (OQ-46): its per-vector MLP makes cross-architecture byte-identical encoding (the 0004 §5.4 determinism gate) an open problem, so it is a candidate algorithm, not a v0.1 conformance-gated one. Depends on: spec/0001, spec/0002, spec/0004, spec/0007. Motivation: spec/0004 v0 ships partitioning algorithms (LSH, IVF, IMI) that decide where a vector lives, but stores every vector as raw binary32. At billion scale this dominates total storage and bucket-fetch latency: 1B × 3 KB = 3 TB of vector bytes, ~12 MB per Spatial Bucket. Production-grade ANN systems (FAISS-IVF-PQ, ScaNN, DiskANN) compress vectors 30–100× via Product Quantization or its neural successors. This spec defines the DreamDB-side primitives — a VectorCompressor Object, modality-string extension, and bucket-layout impact — needed to swap raw f32 for quantized codes without disturbing the partitioning algorithms or the address grammar.

The v0.1 registered + implemented algorithms are dreamdb.raw-f32 (the no-compression baseline), dreamdb.rabitq-cosine (binary quantization with a deterministic rotation — the recall-at-footprint headliner: fast, byte-deterministic, ADC-scored, with optional per-record correction factors and 1/2/4/8-bit variants), and dreamdb.pq-cosine (Product Quantization, the well-understood codebook baseline). dreamdb.qinco-cosine (Huijben et al., NeurIPS 2025), a learned residual quantizer with neural codebooks, is reserved for a future revision (OQ-46): its quality is state-of-the-art at a fixed bitrate, but its per-vector neural-network forward pass is both an ingestion-throughput cost and a cross-architecture determinism hazard — so it is better placed as a performance-profile (0023) advisory algorithm than a conformance-gated normative one. The spec is named for the framework (the VectorCompressor Object + registry of algorithms), not for any single algorithm.


1. Purpose

Vector compression is orthogonal to vector partitioning. spec/0004 answers "which bucket does this vector live in"; spec/0010 answers "what bytes do we store for the vector once it's in the bucket." A modality SHOULD be free to combine dreamdb.imi-cosine partitioning with dreamdb.qinco-cosine compression, or dreamdb.lsh-cosine partitioning with raw f32 storage, without either subsystem leaking knowledge of the other.

By the end of this document the following are concrete:

  • The VectorCompressor contract: signature, determinism, and distributability requirements for any function that encodes f32 vectors to compact codes.
  • The VectorCompressor Object: the immutable, content-addressed CBOR Object that distributes encoder/decoder weights — symmetric with the SpatialIndex Object of spec/0004.
  • The compressor binding: Schema and the modality's registry reference identify the VC Object governing bucket records, without a new modality suffix.
  • The bucket-layout impact: how spec/0007 §6.1 inline records and §6.3 VS Objects change when record_size is determined by code length rather than dim × 4.
  • The re-rank contract: query path semantics for "fetch compressed codes, identify top-K candidates, optionally fetch full-precision vectors for exact distance."
  • The registered algorithms: dreamdb.raw-f32 (no compression — the default for backwards compatibility), dreamdb.rabitq-cosine (binary quantization — the v0.1 recall-at-footprint default), and dreamdb.pq-cosine (classical PQ) are normative and implemented; dreamdb.qinco-cosine (learned RQ, §7) is reserved for a future revision (OQ-46).
  • The neural-determinism discipline: a tightened version of spec/0004 §5.4.1's f32 rules that handles the harder reproducibility surface of neural codebooks.

What stays defined elsewhere:

  • The <spatial-key> derivation (partitioning) — spec/0004.
  • The Bucket Object header and reference-mode VS Object header — spec/0007 §6.1, §6.3.
  • The Manifest registry shape — spec/0002 §7.2.

What this document does NOT define:

  • Query-side ranking algorithms (e.g. ScaNN's anisotropic loss) — those are SDK implementation choices.
  • Adaptive bit-budget policies (decide compression level per-bucket) — deferred to v0.X+1.
  • Joint train of partitioner + compressor (e.g. SPANN-style) — deferred. spec/0010 assumes the two are trained independently.

2. The VectorCompressor Contract

A VectorCompressor is a deterministic, distributable pair of functions that map between vectors and compact byte codes:

encode : Vector_D → Codes_C       (C bytes; C ≪ 4 × D)
decode : Codes_C  → Vector_D      (approximate reconstruction; optional for ANN paths)

Conformant implementations MUST satisfy three properties paralleling spec/0004 §2:

2.1 Encode determinism

encode(v) == encode(v')   whenever v == v' bit-for-bit

Two implementations of the same VectorCompressor Object MUST produce bit-identical codes for identical inputs. No tolerance for hardware FP-determinism drift; if the reference codebook training was done on GPU, the inference path MUST still reproduce the scalar-CPU reference output to the bit. See §8 for the neural-determinism discipline.

2.2 Distributability

The VectorCompressor's full behavior MUST be reproducible from a single immutable Object (the VectorCompressor Object). Codebooks, network weights, normalization stats — everything required to encode (and to decode, if the algorithm provides decode) — lives in this Object. No hidden state.

2.3 Approximation bound

Each algorithm SHOULD declare a reconstruction-error or recall-at-fixed-budget property that an operator can use to size the storage/recall tradeoff. The contract requires only that one be provided in the algorithm's §11 entry; specific bounds are algorithm-specific.

3. The VectorCompressor Object

Address path:

vector-compressor/<multihash-of-canonical-CBOR-bytes>

(New top-level namespace, parallel to spatial-index/ and scalar-index/. Lives outside the per-Timeline tree because a single compressor MAY be shared across many Timelines.)

3.1 CBOR encoding

{
  "algorithm":   "<algorithm-id>",          ;; e.g. "dreamdb.qinco-cosine"
  "dim":         <unsigned int>,            ;; input dimensionality D
  "code_bytes":  <unsigned int>,            ;; output code length C, in bytes
  "metric":      "<metric-id>",             ;; "cosine" or "l2"
  "supports_decode": <bool>,                ;; true ⇒ approximate decode is available
  "params":      <algorithm-specific>,      ;; see §6, §7
}

All fields are required. supports_decode = false means the compressor is encode-only and the SDK MUST issue a re-rank fetch (§5) for any query needing exact distances.

3.2 Reference from the Manifest registry

A modality that uses bucket-internal compression MUST declare its VectorCompressor Object hash in the same registry entry that already declares the SpatialIndex:

"registry": {
  "embedding.f32.dim=768.bucketed": {
    "kind":               "continuous",
    "object_kind":        "spatial-bucket",
    "algorithm":          "dreamdb.imi-cosine",
    "spatial_index":      [<multihash-SI-Object>],
    "vector_compressor":  <multihash-VC-Object>,            ;; NEW (this spec)
    "replicate_probes":   0,
  }
}

vector_compressor is OPTIONAL. Absent or null ⇒ behavior equivalent to dreamdb.raw-f32 (no compression). When present, readers MUST validate per §3.3.

The exact-vector path is per candidate, and there is no registry-level counterpart. An inline SpatialBucketEntry MAY carry a rerank_storage_hash naming a Vector-Storage Object (per spec/0007 §6.3) that holds full-precision vectors. Its records align with that Bucket by ordinal. A compressed reference record instead carries its own (vec_obj_hash, byte_offset) raw location, because one Bucket may point into several VS Objects. Different candidates of one Track may therefore name different exact Objects, or none.

A reader MUST therefore resolve exactness per candidate and group its fetches by the VS Object each one names. Treating a Track as having a single sidecar would read the wrong object for some candidates and miss that others have none — which is also why "does this Track have exact vectors" is not a yes/no property.

Where an inline candidate's entry carries no rerank_storage_hash, or a future record form carries no exact locator, the compressed distance is the only signal available — appropriate for "cheap recall" use cases (training data filtering, coarse semantic search) where exact distance is not required.

3.3 Registry-vs-VC consistency (mandatory validation)

The SDK MUST validate at load time:

  1. Compressor resolution: vector_compressor is an Object hash, not an algorithm-name field. Resolve it to the VectorCompressor Object and check that Object's algorithm is supported; the registry's algorithm names the independent partitioner and need not equal it.
  2. Dimensionality: VC Object's dim equals the modality's dim parameter and equals the SpatialIndex Object's dim.
  3. Code width: the algorithm's parameters and record layout must agree with the VC Object's code_bytes. The registry binding (§4), not a tag suffix, supplies the compressor identity.
  4. Bucket-header agreement: per spec/0007 §6.1.1, every Bucket fetched MUST carry a header field vector_compressor_hash matching the VC Object the SDK is currently using. Mismatch ⇒ critical error per §6.1.1 (this spec extends the lineage-validation discipline to compressors).

A contradictory binding or record layout is malformed; an otherwise coherent but unsupported algorithm is a capability refusal, not proof of corruption. Both must stop the affected decode/query rather than interpret codes using the wrong codebook. The lineage risk is the same as for spatial_index_hash (0004).

4. Compressor binding; no modality-string extension

Compression is declared by the modality's vector_compressor registry reference (§3.2), consistent with its field's Schema compressor binding. The referenced VC Object supplies algorithm, code width and parameters. A reader MUST resolve that binding and enforce lineage (§3.3); absence of a compress= tag suffix MUST NOT be interpreted as absence of compression. For example, embedding.f32.dim=768.bucketed can name a compressed Track when its registry entry supplies the VC reference.

The former .compress=qinco:M=8 proposal is withdrawn, not replaced with a new suffix. Its colon and second equals sign violate 0002 §5.1; uppercase letters in a parameter value are not themselves invalid. No normalization, alias or rewrite of existing tags is authorized. This retains the reference writer's existing Schema/registry contract and does not add a second parameter source.

Bucket header size is selected by the encoded header_size, not inferred from tag spelling: 160 for the original layout, 200 for the layout carrying a VC hash (0007 §6.1.2). Header lineage and record width must agree with the resolved binding. An unsupported compressor or header is refused, not decoded as raw f32.

5. Re-rank contract (query path)

When vector_compressor is set on a modality, the candidate-fetch step of the query verb (per spec/0006 §6.6, post-spec/0004 §6.5) becomes:

1. Compute spatial keys (per spec/0004) for the query vector q.
2. Fetch candidate Bucket Objects (per spec/0004 §6.5).
3. For each Bucket:
   - Decode the COMPRESSED code for each record (cheap; ~µs per record).
   - Compute ADC (asymmetric distance) using a per-query lookup table
     against the codebooks held in the VC Object.
4. Rank by compressed-distance; keep top K_compressed candidates
   (K_compressed ≥ K, typically 4× to 8×).
5. IF the effective query policy requires exact ranking (Schema by default):
   - Resolve each candidate's exact source: an inline entry's
     `rerank_storage_hash` plus record ordinal, a reference record's raw VS
     hash plus byte offset, or the raw bytes already held by a HotShard.
   - Group remote sources by VS Object and issue one multi-range GET per
     distinct Object, covering that group's record byte ranges.
   - Compute exact f32 distance.
   - Re-sort; return top K.
   ELSE:
   - Return top K by compressed-distance directly.

Steps 1–4 are the production hot path. Step 5 is decided by the field's Schema — operators trade an additional GET round-trip and ~K×4D byte fetch for exact distances.

What decides, and over what. The Schema's rerank flag supplies the default; only an explicit query mode may override it. Data availability never chooses the policy. Opportunistically reranking whichever records have sidecars would mix exact and approximate scores in one ordering and is forbidden.

A query implementation MUST therefore, when inheriting the Schema policy (this is the Dataset layer's obligation, not the Connector's — the Connector is pure transport per spec/0005 §8.4):

  • with rerank: false, order approximately throughout and read no rerank storage, whatever references exist;
  • with rerank: true, classify exact-vector coverage over the selected rerank pool after eligibility filtering — the candidates that remain once tombstones and scalar predicates have been applied, and that the pool has admitted — and refuse the query unless every one of them has an exact source;
  • return Ok([]) when nothing survives filtering: an ordinary query miss is not a coverage failure.

The pool, not the whole filtered candidate set, is the scope because the pool is what can be returned: nothing outside it reaches the result, so a missing exact source beyond it cannot affect the query's correctness, while any member may be promoted into the result by exact scoring and so must have one. The pool is normally wider than top_k, so checking only the current approximate top_k would be too narrow.

Eligibility is applied before coverage is classified and before the pool is selected. A record a tombstone or a where clause removes is not a candidate: it must not make a covered field look partially covered, and must not occupy a pool slot that a returnable record could hold.

Candidates read from a HotShard buffer (spec/0016 §2.5) hold raw f32 and are exact-capable, though they carry no rerank_storage_hash; a query implementation MUST NOT infer their exactness from that field alone. Where a compressor is configured, hot and cold candidates MUST be scored on the same approximate measure for pool selection — comparing an exact cosine against an ADC estimate in one ordering is not a ranking.

Pool sizing is a separate, query-local tuning option. The Dataset partition-query API accepts VectorQueryOptions.rerank_pool_size (Python rerank_pool_size, JavaScript rerankPoolSize). Absence preserves the historical saturating 5 * top_k default. An explicit size MUST be positive and at least top_k, including when top_k is zero; invalid values are configuration errors. Rust/Python sizes must fit the target's usize; JavaScript accepts only integer Numbers in [1, 2^32-1], never truncating fractions, negatives or overflow. After eligibility, the effective pool is min(requested_size, candidate_count); the request does not preallocate that many candidates or expand ANN probing. Ordinary and identity-aware queries use the same rule. Non-vector and graph queries MUST reject an explicit pool rather than silently ignore it. The latter have separate traversal/reranking contracts, not this partition pool.

Pool tuning does not alter the Schema's exactness requirement, permit partial exact coverage, or change the interpretation of returned scores. A pool multiplier is not a recall or latency guarantee. Exact rescoring cannot recover a candidate excluded from the probe result or selected pool; measurements require a pinned corpus, query set and environment (0023).

Explicit modes (OQ-44 resolved). VectorQueryOptions.rerank_mode (Python rerank_mode, JS rerankMode) accepts exactly inherit, approximate, or exact. Absence means inherit and preserves the Schema rule above. Explicit approximate keeps selection scores and MUST NOT read rerank storage, even for a Schema that defaults to exactness. This is an explicit caller choice, not an automatic fallback, and changes neither stored Schema nor the interpretation of the metric. Explicit exact requires exact scores for every selected member; missing capability is a configuration refusal when an approximate Schema is healthy but cannot satisfy the request, and a Protocol refusal when published data violates the Schema's own exactness declaration. There is no partial fallback. An empty selected pool returns an empty result in all modes.

Raw, uncompressed candidates already hold exact metric scores; explicit exact does not demand redundant sidecars there. Compressed approximate scores remain ADC estimates; exact cosine scores come from raw vectors, and raw L2 scores are negative squared Euclidean distances. No mode guarantees global ANN recall. Graph/non-vector queries reject non-inherited modes; graph traversal continues to use its existing contract. This resolves override semantics without adding graph tuning, persisted fields, or a new format discriminator.

5.1 ADC lookup-table construction

For Product Quantization and QINCo, the per-query ADC table is constructed once at step 1:

For each sub-codebook s ∈ {0, …, M-1}:
  For each centroid c ∈ {0, …, K_centroids-1}:
    LUT[s][c] = <q_subvector_s, codebook[s][c]>

In-bucket distance for a record with codes (c_0, …, c_{M-1}):

distance(q, record) = Σ_{s=0}^{M-1} LUT[s][c_s]

M × K_centroids floating-point multiply-adds at table construction; M table lookups + adds per record at scoring. With M=8, K_centroids=256, this is 2048 multiplies upfront then 8 lookups per record — ~100× faster than computing <q, decoded_v> per record.

5.2 Re-rank storage layout

Every exact-source VS Object follows the existing spec/0007 §6.3 format unchanged. Inline mode names a parallel Object from SpatialBucketEntry.rerank_storage_hash; its records align with that Bucket by ordinal. Reference mode names the raw Object and absolute byte offset in each record (§8.3), so no one-to-one entry-level sidecar is implied. In both modes, compressed codes live in Bucket Objects and exact f32 records live in the VS Objects the candidate identifies.

Operators with strict storage budgets MAY write bucket entries that carry no rerank_storage_hash at all — compressed-only mode. The recall gap from re-rank is then unrecoverable, but storage drops to code_bytes/4D of the uncompressed baseline. Acceptable for many ML-training and indexing workloads.

6. v0.X Default Algorithm: dreamdb.raw-f32

The trivial compressor. Provided so the algorithm registry covers the "no compression" case symmetrically — Manifests without a vector_compressor field behave identically to Manifests declaring algorithm: "dreamdb.raw-f32".

;; params
{
  "version": 1,
}

code_bytes = 4 × dim. supports_decode = true (encode and decode are both the identity). No training data required.

7. v0.X Algorithm: dreamdb.qinco-cosine

dreamdb.qinco-cosine is a learned residual quantizer with neural codebooks (Huijben et al., NeurIPS 2025). Compression quality at billion-scale dim=768: at M=8 codes (8-byte records ≈ 384× compression), recall@10 within ~2 percentage points of exact f32 on the BIGANN-1B benchmark; classical PQ at the same budget loses ~10 pp.

The structural difference from classical PQ:

  • Classical PQ uses a fixed codebook per sub-quantizer, learned once. Quantization is independent across sub-quantizers.
  • QINCo uses a small MLP that produces codebook deltas conditioned on the residual prefix. Each sub-quantizer's effective codebook adapts to the partial reconstruction so far. The MLP weights are the algorithm's identity.

7.1 Params (CBOR)

{
  "version":         1,
  "M":               <uint>,                ;; number of residual stages (= code_bytes)
  "K":               <uint>,                ;; centroids per stage (typically 256)
  "stage_codebooks": <byte-string>,         ;; M × K × dim LE-f32 row-major
  "mlp_layers":      [                      ;; per-stage MLP weights (small)
    { "w0": <bytes>, "b0": <bytes>,
      "w1": <bytes>, "b1": <bytes> },
    …                                       ;; one entry per stage
  ],
  "training_metadata": {                    ;; OPTIONAL; reproducibility manifest
    "dataset_hash":  <bytes>,
    "epochs":        <uint>,
    "lr":            <float-as-bytes>,
    "seed":          <bytes>,
  },
}

code_bytes = M (one byte per stage when K=256; the registered grammar permits other K, with proportional encoding).

7.2 Encoder pseudocode

encode(v):
  residual ← v
  codes    ← []
  for stage s in 0..M:
    cb        ← stage_codebooks[s]            ; (K, dim)
    delta     ← mlp_s(residual)               ; (K, dim) adaptive deltas
    effective ← cb + delta
    j         ← argmax_{i ∈ 0..K} <residual, effective[i]>
    codes.append(j)
    residual  ← residual − effective[j]
  return codes

Decode is the reverse: reconstruct effective per stage, lookup the centroid, accumulate. supports_decode = true.

7.3 Floating-point and neural determinism

QINCo inherits all of spec/0004 §5.4.1's f32 discipline (no FMA, no -ffast-math, scalar reference path) and adds:

7.3.1 MLP inference determinism (mandatory)

The MLP's forward pass MUST run on a fixed-order scalar reference path in the conformance test suite. Implementations MAY use SIMD or accelerator backends for performance but MUST verify against the scalar reference on every supported architecture (per 0009 §5.3.1).

Practical implication: a QINCo VC Object trained on GPU MUST round-trip through CPU inference and produce bit-identical codes for the conformance test vectors. This is achievable in practice — modern frameworks (ONNX Runtime in CPU-determinism mode, PyTorch with torch.use_deterministic_algorithms(True) + CPU backend) can match the scalar reference if compiled without fast-math and run on a single thread.

Producers SHOULD treat training as a separate phase from publishing: train on whatever hardware is convenient, then re-encode the codebooks on the scalar reference path before producing the VC Object's stage_codebooks bytes. The conformance suite ships round-trip vectors (v → encode(v) → codes) for both K=256, M=8, dim=768 and K=256, M=16, dim=128; an implementation that disagrees on any test vector is non-conformant.

7.3.2 Activation function choice

For determinism, QINCo's MLP MUST use ReLU activations, not GELU or SiLU. ReLU is a piecewise-linear function with exact f32 representation; GELU/SiLU require transcendental approximations whose bit-level result varies by library version (libm vs. mlibc vs. musl etc.). The training reference also uses ReLU; this is not a performance-vs-quality tradeoff but a determinism requirement.

7.4 Training and the codebook-publication contract

Same shape as spec/0004 §5.6.5 (IVF) and §5.7.5 (IMI):

  1. One writer trains and publishes the VC Object. Its content hash is the compressor's identity on the timeline.
  2. Subsequent writers fetch and consume. The Manifest's vector_compressor field carries the hash; downstream writers encode against the published codebooks, NOT re-trained ones.
  3. No mid-Track re-training. A second-generation codebook goes into a new VC Object referenced from a new Track (per spec/0008 Track layering). Mutating an existing VC Object is FORBIDDEN.

Training procedure is RECOMMENDED to follow the QINCo reference codebase (Huijben et al. 2025, Algorithm 1) with the determinism harnesses of §7.3. Implementations that train independently MUST agree on:

  • Training-data sample (canonical: first N vectors of the dataset, with N declared in training_metadata).
  • RNG seed.
  • Optimizer (AdamW), learning rate, schedule, epoch count — all declared in training_metadata.

training_metadata is OPTIONAL because operators may legitimately publish a VC Object without disclosing their training pipeline; the field is for reproducibility, not for protocol-level validation.

8. Bucket-layout impact (cross-reference to spec/0007 §6)

This spec amends spec/0007 §6.1 (inline-mode Bucket layout) and §6.1.1 (lineage validation):

8.1 Record size becomes code-byte size

Per-record layout when vector_compressor is non-null:

record_size = 8 (time_anchor) + code_bytes

code_bytes comes from the VC Object, not from the modality's dim parameter. For a hypothetical 8-byte code, records are 16 bytes — about 192× smaller than a dim=768 f32 record (3080 bytes); this is not a claim that QINCo ships. Record offsets are header_size + i × (8 + code_bytes); the code itself starts eight bytes later, after the anchor. A VC-bearing inline Bucket has header_size=200, not 160 (0007 §6.1.2). Modality parameters alone do not determine these offsets.

8.2 Bucket header carries vector_compressor_hash

The compressed Bucket uses the 200-byte header defined by spec/0007 §6.1.2. Its byte layout is:

0..4      magic
4..8      version
8..12     record_size
12..16    record_count
16..20    header_size = 200
20..53    spatial_index_hash
53..85    modality (32-byte zero-padded ASCII prefix)
85..86    bucket mode
86..119   vector_compressor_hash
119..200  reserved (81 zero bytes)

The vector_compressor_hash begins at offset 86, after the modality and mode byte. Writers MUST zero-fill 119..200; readers MUST ignore those reserved bytes. The reservation is retained for future compatible extensions and contains no magic suffix.

Legacy v0 Buckets have a 160-byte header and do not contain a compressor-hash field. Decoders determine the header shape from the modality and the encoded header_size; they MUST NOT reinterpret the zero-filled 86..119 portion of a 160-byte header as a compressor hash. A 200-byte compressed header carries a real multihash at 86..119; an all-zero value there is malformed, not an alternate spelling of raw-f32.

8.3 Reference-mode interaction (spec/0007 §6.2)

Reference mode combined with compression uses inline codes beside the raw reference. The header is the 200-byte form from §8.2. Each record is:

time_anchor:u64be | vec_obj_hash:multihash(33) | byte_offset:u64be | code:code_bytes

Thus record_size = 49 + code_bytes. The code is scored during candidate selection without an additional lookup. The first 49 bytes retain the raw VS location used by exact reranking. Writers and readers MUST preserve the association within each record; compaction may reorder or regroup records only as whole units and MUST preserve both the compressor lineage and code width.

Codes are not stored in a parallel VS segment. That alternative would require another ordinal-to-code mapping, another fetch, and another closure rule while the reference record is already fetched for membership. The selected layout therefore resolves OQ-45 without adding an Object address or changing the raw Vector-Storage format.

9. Compression vs. partitioning composition

Algorithms register independently. Manifests declare:

  • algorithm (partitioner): dreamdb.lsh-cosine | dreamdb.ivf-cosine | dreamdb.imi-cosine | future.
  • vector_compressor (compressor): dreamdb.raw-f32 | dreamdb.rabitq-cosine | dreamdb.pq-cosine; QINCo remains reserved, not a supported-algorithm requirement.

The framework separates partitioning from compression. A supported combination must satisfy dimension, metric, record-shape and lineage constraints on both sides; this is not a claim that every implementation supports every Cartesian product, or that reserved algorithms are conformant. Reference mode uses §8.3. Unsupported combinations must be reported explicitly, not decoded as raw vectors.

                          ┌──────────────────┐
                          │   query vector q │
                          └────────┬─────────┘
                                   ├────────────────────────────────────┐
                                   ▼                                    ▼
                       ┌─────────────────────┐         ┌─────────────────────────────────┐
                       │ partitioner.encode  │         │ compressor.encode  (codebook)   │
                       │  → spatial_key      │         │ AND                              │
                       │  (spec/0004)        │         │ compressor.build_LUT(q)         │
                       └──────────┬──────────┘         │ for ADC scoring (spec/0010 §5.1)│
                                  │                    └─────────────────────┬───────────┘
                                  │                                          │
                                  ▼                                          ▼
                       (fetch buckets at spatial_key)     (score in-bucket records via LUT)

The two pipelines join only at "candidate Bucket fetch → compressed-record scoring." Neither knows the other's internals.

10. Storage and latency at billion-scale

Informative payload arithmetic, not a measured storage/latency profile. Assume 1B vectors, dim=768, 3,800 records per fetched Bucket and 16 Buckets per query. The table counts vector/code bytes only: add anchors, headers, indexes, table replication and retained raw exact sources. These occupancy assumptions are not derived from a particular IMI parameterization. QINCo rows illustrate a proposed code width, not a shipped capability.

Compressorcode_bytesTotal vector storagePer-Bucket size (~3800 records)Wire fetch per query (~16 buckets)Re-rank fetch (K=10, K_comp=40)
dreamdb.raw-f3230723.0 TB~12 MB~192 MBn/a (already exact)
dreamdb.pq-cosine M=888.0 GB~32 KB~512 KB~120 KB (40 × 3072 bytes)
dreamdb.qinco-cosine M=888.0 GB~32 KB~512 KB~120 KB (40 × 3072 bytes)
dreamdb.qinco-cosine M=444.0 GB~16 KB~256 KB~120 KB

An 8-byte code is 384 times smaller than a 3,072-byte vector before overhead. This ratio is neither total storage reduction nor measured query speedup. In particular, raw-f32 gives exact scores within the visited candidates, not recall@10 = 1 for an ANN query. This specification supplies no measured recall or latency result for the table; such a result belongs in a reproducible 0023 profile report.

11. Algorithm registry additions

Extends spec/0004 §3.4:

Algorithm IDFamilyDefined inRecommended use
dreamdb.raw-f32compressor§6Default; backwards-compatible with v0
dreamdb.rabitq-cosinecompressorFramework / 0009 §8.5.1Implemented binary-quantization family
dreamdb.pq-cosinecompressor§11.1 / 0009 §8.5.1Implemented codebook baseline
dreamdb.qinco-cosinereserved compressor§7 proposalDeferred pending OQ-46; not a production conformance claim

11.1 dreamdb.pq-cosine (implemented baseline)

When paired with dreamdb.ivf-cosine, this remains a compressor of the original global vector, not an IVF residual codec. 0004 §5.6.1 defines independent training/binding, probing, ADC and exact-rerank composition. Neither a new pq-ivf identity nor per-cell codebooks are introduced.

Classical Product Quantization. The reference PqCosineParams carries version, m, k, codebooks and training_seed; D is divisible by m, and each of the m subspaces has k centroids, with 1 ≤ k ≤ 256. Codebooks are f32le bytes in subspace/centroid/dimension order. One byte per subspace selects a centroid, so code_bytes = m. Encoding maximizes the sub-vector/centroid dot product using the scalar f32 accumulation discipline, with the smaller centroid index winning ties; decoding concatenates the chosen centroids. The format and reference behavior are exercised under 0009 §8.5.1. This is an implemented algorithm, not permission to claim a new user-defined spelling. Comparative recall and throughput are workload-dependent and are not specified here.

12. Out of scope

  • Per-bucket adaptive compression. Future: code budget varies by bucket density / query frequency. Defers to v0.X+1.
  • Joint train of partitioner + compressor. SPANN-style integrated training MAY land as a single combined algorithm ID (e.g. dreamdb.spann-cosine) without reusing this spec's separation. Deferred.
  • Cross-modality codebook sharing. A VC Object is keyed by dim and metric only; nothing in the protocol prevents two modalities with the same dim from referencing the same VC Object. Address-space dedup happens automatically. No special protocol affordance needed.
  • Non-cosine metrics. dreamdb.qinco-l2 etc. follow the same template; defer until L2-metric modalities ship per spec/0004 OQ-26.

13. Open questions

  • OQ-44 (→ spec/0006): Resolved. §5 defines query-local pool sizing and explicit inherit/approximate/exact modes for partition queries, with strict exact-capability refusal and unchanged default behavior. Graph overrides are not supported and explicitly refuse.
  • OQ-45 (→ spec/0007): Resolved. Reference-mode records inline one fixed-width compressed code after their existing 49-byte raw-vector reference (§8.3). The code participates in approximate ranking; the reference remains the exact-rerank source. A parallel codes segment was rejected because it adds lookup, mapping, and reachability machinery without removing the already-required Bucket read.
  • OQ-46 (→ spec/0009): Deterministic encoding and portable vectors for dreamdb.qinco-cosine, including a pinned small reference VC Object. This blocks promotion of QINCo to a normative supported algorithm, not releases of the already-supported compressors.
  • OQ-47 (→ this spec): Resolved. The implementation and already-written compressed Buckets are authoritative for the compatible layout: vector_compressor_hash occupies 86..119, followed by 81 reserved bytes at 119..200. Writers zero-fill the reservation and readers ignore it. No magic suffix is added because doing so would change the content address of existing Buckets.

Next: spec/0010 amendments after first implementation pass (likely tighter determinism guidance and concrete conformance vectors).