DreamDB

Indexes and compressors

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

DreamDB stores a vector by where it belongs, so the thing that decides where has to exist before the first write. Training it is one call, and the result is a content hash you pass into the schema.

Partitioning indexes

Cosine partitioning fields require a published index, for example:

python
index = db.train_and_publish_ivf_centroids(
    BACKEND, sample, dim=32, k=16,       # sample: (N, dim) float32
)

schema = db.Schema().add_embedding(
    "embedding", dim=32, algorithm="dreamdb.ivf-cosine", spatial_index=index,
)
AlgorithmTrain withChoose when
dreamdb.ivf-cosinetrain_and_publish_ivf_centroids(backend, sample, dim, k)The default choice. k ≈ √N for the production corpus
dreamdb.imi-cosinetrain_and_publish_imi_centroids(backend, sample, dim, k_sub)Very large corpora, where a flat centroid list gets unwieldy
dreamdb.lsh-cosinePublished hyperplanesCosine partitioning; Node authoring can construct this path
dreamdb.lsh-l2L2 Schema pathMagnitude-preserving Euclidean search; see L2 examples

Training is deterministic — the same (sample, dim, k, iterations, seed) yields the same 33-byte hash. Two workers that train independently publish the same object, which is a content-addressed no-op rather than a conflict.

Sizing the sample matters more than sizing the corpus. Ten to a hundred thousand vectors sampled from real data trains a better index than a million synthetic ones. Operators typically extract the sample offline and keep it, so the index can be rebuilt reproducibly.

Compressors

A compressor changes what is stored per vector, not where it goes. Pair one with any partitioning algorithm:

python
compressor = db.publish_rabitq_compressor(BACKEND, dim=512, bits_per_dim=1)

schema = db.Schema().add_embedding(
    "clip", dim=512, algorithm="dreamdb.ivf-cosine",
    spatial_index=index, compressor=compressor, rerank=True,
)
CompressorPublish withNotes
dreamdb.raw-f32Nothing to publish — the defaultFull precision, largest footprint
dreamdb.rabitq-cosinepublish_rabitq_compressor(backend, dim, bits_per_dim=1)Data-oblivious, so no training data needed. 1/2/4/8-bit variants
dreamdb.pq-cosinetrain_and_publish_pq_compressor(backend, sample, dim, m, k)Classical product quantization; needs a training sample

rerank=True keeps full-precision vectors alongside the compressed codes, so a query scores candidates cheaply and then re-scores the survivors exactly. It costs storage and buys back the accuracy compression gave away.

dreamdb.qinco-cosine appears in the specification as reserved but deferred — its per-vector neural forward pass makes byte-identical cross-architecture encoding an open problem. See Vector Compression.

Graph indexes

For graph-based ANN over a field that already exists:

python
db.train_and_publish_vamana_graph(BACKEND, sample, dim=128, modality="…",
                                  r=32, alpha=1.2, l_build=100)

r is the out-degree, alpha the pruning slack, l_build the build-time beam width. Larger values buy recall with build time and index size. See Graph Indexing.

Text indexes

BM25 over a text field, built from (anchor, text) pairs. No model is involved — the tokenizer is built in:

python
anchors = ds.list_anchors()
docs = [(a, caption_for(a)) for a in anchors]

ds.add_text_index("caption_idx", "caption", docs)     # k1=1.5, b=0.75 by default

ds.query_text("caption_idx", "dog", 3)
# [(1785381710580097000, 1.2619806528091431), …]

k1 and b are the usual BM25 knobs. Positions are off by default; turn them on only if you need phrase queries, since they enlarge the index.

A dataset carrying a text index cannot currently be union-merged with a branch, because the text modality's index lineage diverges between trunk and branch. The refusal is deliberate and explicit — spec/0008 §6.5 requires it rather than silently mixing incompatible indexes — but it means text indexes and sharded ingest do not yet combine. Build the text index after merging.

Keeping storage in shape

python
ds.compact()
# {'cells_examined': 21, 'cells_compacted': 0, 'fragments_collapsed': 0, …}

db.gc(BACKEND, keep_manifests=10, dry_run=True)

Bulk ingest leaves many small fragments per cell, which turns one logical read into many object GETs. compact() merges them. gc removes unreferenced objects and defaults to dry_run=True — read what it proposes before letting it delete.