DreamDB

Quickstart

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

Forty records with embeddings, labels, and captions, written to a directory and queried back. No services, no cloud account.

bash
pip install dreamdb numpy
mkdir -p /tmp/notes

Train, declare, create

A partitioning index decides where each vector is stored, so it has to exist before the first record does. dreamdb.ivf-cosine needs a sample of representative vectors to train on — a few hundred is enough to start, and k ≈ √N is the rule of thumb for the production corpus size:

python
import numpy as np
import dreamdb as db

BACKEND = "file:///tmp/notes"
rng = np.random.default_rng(5)

index = db.train_and_publish_ivf_centroids(
    BACKEND, rng.standard_normal((512, 32)).astype(np.float32), dim=32, k=16,
)

schema = (db.Schema()
          .add_embedding("embedding", dim=32,
                         algorithm="dreamdb.ivf-cosine", spatial_index=index)
          .add_scalar_categorical("label")
          .add_scalar_string("caption"))

ds = db.Dataset.create("notes", schema, backend=BACKEND)

Append

One dict per record, keyed by field name. numpy arrays are accepted directly:

python
captions = ["a brown dog on grass", "a tabby cat asleep",
            "a red car at night", "two dogs running"]

ds.append_many([
    {"embedding": rng.standard_normal(32).astype(np.float32),
     "label": ["cat", "dog"][i % 2],
     "caption": captions[i % 4]}
    for i in range(40)
])

print(ds.count())                      # 40
print(ds.distinct_values("label"))     # [('cat', 20), ('dog', 20)]

append_many commits by default. Pass commit=False to batch several appends into one manifest and call commit() when you are done.

Query

Vector search returns batches, not rows — the reader streams whole buckets:

python
query = rng.standard_normal(32).astype(np.float32)

for batch in ds.iter_vector(field="embedding", query=query, top_k=5, batch_size=5):
    print(batch["label"], batch["_time_anchors"][:1])

Every batch is a dict of columns plus _time_anchors. Add as_numpy=True to get the embedding column as an ndarray instead of a list.

Filters compose with the vector search, so the scan never widens:

python
hits = ds.iter_vector(field="embedding", query=query, top_k=5,
                      where_eq={"label": "dog"}, as_numpy=True)

Scalar-only iteration skips vectors entirely:

python
rows = sum(len(b["_time_anchors"]) for b in ds.iter_scalar(where_eq={"label": "cat"}))
print(rows)     # 20

Reopen it later

The schema is recovered from the manifest, so it does not need re-declaring:

python
ds = db.Dataset.open("notes", backend=BACKEND)

Full-text search

Captions are searchable once a BM25 index exists over them. Build it from (anchor, text) pairs:

python
anchors = ds.list_anchors()
docs = [(a, captions[i % 4]) for i, a in enumerate(anchors)]

ds.add_text_index("caption_idx", "caption", docs)

print(ds.query_text("caption_idx", "dog", 3))
# [(1785381710580097000, 1.2619806528091431),
#  (1785381710580101000, 1.2619806528091431),
#  (1785381710580105000, 1.2619806528091431)]

query_text returns (anchor, score) tuples. The tokenizer is built in, so no model is involved — more on indexes.

Next