# Indexes and compressors

> **Version scope:** Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive
> geometry, structured ragged/CSR and additional index capabilities are included
> in these releases. See [feature examples and limits](/main-features.md) for the
> supported boundaries; older packages may not expose these methods.

DreamDB stores a vector by *where it belongs*, so the thing that decides where has to exist before the first write. Training it is one call, and the result is a content hash you pass into the schema.

## Partitioning indexes

Cosine partitioning fields require a published index, for example:

```python
index = db.train_and_publish_ivf_centroids(
    BACKEND, sample, dim=32, k=16,       # sample: (N, dim) float32
)

schema = db.Schema().add_embedding(
    "embedding", dim=32, algorithm="dreamdb.ivf-cosine", spatial_index=index,
)
```

| Algorithm | Train with | Choose when |
|---|---|---|
| `dreamdb.ivf-cosine` | `train_and_publish_ivf_centroids(backend, sample, dim, k)` | The default choice. `k ≈ √N` for the production corpus |
| `dreamdb.imi-cosine` | `train_and_publish_imi_centroids(backend, sample, dim, k_sub)` | Very large corpora, where a flat centroid list gets unwieldy |
| `dreamdb.lsh-cosine` | Published hyperplanes | Cosine partitioning; Node authoring can construct this path |
| `dreamdb.lsh-l2` | L2 Schema path | Magnitude-preserving Euclidean search; see [L2 examples](/main-features.md) |

Training is deterministic — the same `(sample, dim, k, iterations, seed)` yields the same 33-byte hash. Two workers that train independently publish the same object, which is a content-addressed no-op rather than a conflict.

> **Note:** Sizing the sample matters more than sizing the corpus. Ten to a hundred thousand vectors sampled from real data trains a better index than a million synthetic ones. Operators typically extract the sample offline and keep it, so the index can be rebuilt reproducibly.

## Compressors

A compressor changes what is stored per vector, not where it goes. Pair one with any partitioning algorithm:

```python
compressor = db.publish_rabitq_compressor(BACKEND, dim=512, bits_per_dim=1)

schema = db.Schema().add_embedding(
    "clip", dim=512, algorithm="dreamdb.ivf-cosine",
    spatial_index=index, compressor=compressor, rerank=True,
)
```

| Compressor | Publish with | Notes |
|---|---|---|
| `dreamdb.raw-f32` | Nothing to publish — the default | Full precision, largest footprint |
| `dreamdb.rabitq-cosine` | `publish_rabitq_compressor(backend, dim, bits_per_dim=1)` | Data-oblivious, so no training data needed. 1/2/4/8-bit variants |
| `dreamdb.pq-cosine` | `train_and_publish_pq_compressor(backend, sample, dim, m, k)` | Classical product quantization; needs a training sample |

`rerank=True` keeps full-precision vectors alongside the compressed codes, so a query scores candidates cheaply and then re-scores the survivors exactly. It costs storage and buys back the accuracy compression gave away.

`dreamdb.qinco-cosine` appears in the specification as reserved but deferred — its per-vector neural forward pass makes byte-identical cross-architecture encoding an open problem. See [Vector Compression](/spec-vector-compression.md).

## Graph indexes

For graph-based ANN over a field that already exists:

```python
db.train_and_publish_vamana_graph(BACKEND, sample, dim=128, modality="…",
                                  r=32, alpha=1.2, l_build=100)
```

`r` is the out-degree, `alpha` the pruning slack, `l_build` the build-time beam width. Larger values buy recall with build time and index size. See [Graph Indexing](/spec-graph-indexing.md).

## Text indexes

BM25 over a text field, built from `(anchor, text)` pairs. No model is involved — the tokenizer is built in:

```python
anchors = ds.list_anchors()
docs = [(a, caption_for(a)) for a in anchors]

ds.add_text_index("caption_idx", "caption", docs)     # k1=1.5, b=0.75 by default

ds.query_text("caption_idx", "dog", 3)
# [(1785381710580097000, 1.2619806528091431), …]
```

`k1` and `b` are the usual BM25 knobs. Positions are off by default; turn them on only if you need phrase queries, since they enlarge the index.

> **Warning:** A dataset carrying a text index cannot currently be union-merged with a branch, because the text modality's index lineage diverges between trunk and branch. The refusal is deliberate and explicit — `spec/0008 §6.5` requires it rather than silently mixing incompatible indexes — but it means text indexes and [sharded ingest](/python-sdk-versioning.md#sharded-ingest) do not yet combine. Build the text index after merging.

## Keeping storage in shape

```python
ds.compact()
# {'cells_examined': 21, 'cells_compacted': 0, 'fragments_collapsed': 0, …}

db.gc(BACKEND, keep_manifests=10, dry_run=True)
```

Bulk ingest leaves many small fragments per cell, which turns one logical read into many object GETs. `compact()` merges them. `gc` removes unreferenced objects and defaults to `dry_run=True` — read what it proposes before letting it delete.
