Indexes and compressors
Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.
DreamDB stores a vector by where it belongs, so the thing that decides where has to exist before the first write. Training it is one call, and the result is a content hash you pass into the schema.
Partitioning indexes
Cosine partitioning fields require a published index, for example:
| Algorithm | Train with | Choose when |
|---|---|---|
dreamdb.ivf-cosine | train_and_publish_ivf_centroids(backend, sample, dim, k) | The default choice. k ≈ √N for the production corpus |
dreamdb.imi-cosine | train_and_publish_imi_centroids(backend, sample, dim, k_sub) | Very large corpora, where a flat centroid list gets unwieldy |
dreamdb.lsh-cosine | Published hyperplanes | Cosine partitioning; Node authoring can construct this path |
dreamdb.lsh-l2 | L2 Schema path | Magnitude-preserving Euclidean search; see L2 examples |
Training is deterministic — the same (sample, dim, k, iterations, seed) yields the same 33-byte hash. Two workers that train independently publish the same object, which is a content-addressed no-op rather than a conflict.
Sizing the sample matters more than sizing the corpus. Ten to a hundred thousand vectors sampled from real data trains a better index than a million synthetic ones. Operators typically extract the sample offline and keep it, so the index can be rebuilt reproducibly.
Compressors
A compressor changes what is stored per vector, not where it goes. Pair one with any partitioning algorithm:
| Compressor | Publish with | Notes |
|---|---|---|
dreamdb.raw-f32 | Nothing to publish — the default | Full precision, largest footprint |
dreamdb.rabitq-cosine | publish_rabitq_compressor(backend, dim, bits_per_dim=1) | Data-oblivious, so no training data needed. 1/2/4/8-bit variants |
dreamdb.pq-cosine | train_and_publish_pq_compressor(backend, sample, dim, m, k) | Classical product quantization; needs a training sample |
rerank=True keeps full-precision vectors alongside the compressed codes, so a query scores candidates cheaply and then re-scores the survivors exactly. It costs storage and buys back the accuracy compression gave away.
dreamdb.qinco-cosine appears in the specification as reserved but deferred — its per-vector neural forward pass makes byte-identical cross-architecture encoding an open problem. See Vector Compression.
Graph indexes
For graph-based ANN over a field that already exists:
r is the out-degree, alpha the pruning slack, l_build the build-time beam width. Larger values buy recall with build time and index size. See Graph Indexing.
Text indexes
BM25 over a text field, built from (anchor, text) pairs. No model is involved — the tokenizer is built in:
k1 and b are the usual BM25 knobs. Positions are off by default; turn them on only if you need phrase queries, since they enlarge the index.
A dataset carrying a text index cannot currently be union-merged with a branch, because the text modality's index lineage diverges between trunk and branch. The refusal is deliberate and explicit — spec/0008 §6.5 requires it rather than silently mixing incompatible indexes — but it means text indexes and sharded ingest do not yet combine. Build the text index after merging.
Keeping storage in shape
Bulk ingest leaves many small fragments per cell, which turns one logical read into many object GETs. compact() merges them. gc removes unreferenced objects and defaults to dry_run=True — read what it proposes before letting it delete.