Quickstart
Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.
Forty records with embeddings, labels, and captions, written to a directory and queried back. No services, no cloud account.
Train, declare, create
A partitioning index decides where each vector is stored, so it has to exist before the first record does. dreamdb.ivf-cosine needs a sample of representative vectors to train on — a few hundred is enough to start, and k ≈ √N is the rule of thumb for the production corpus size:
Append
One dict per record, keyed by field name. numpy arrays are accepted directly:
append_many commits by default. Pass commit=False to batch several appends into one manifest and call commit() when you are done.
Query
Vector search returns batches, not rows — the reader streams whole buckets:
Every batch is a dict of columns plus _time_anchors. Add as_numpy=True to get the embedding column as an ndarray instead of a list.
Filters compose with the vector search, so the scan never widens:
Scalar-only iteration skips vectors entirely:
Reopen it later
The schema is recovered from the manifest, so it does not need re-declaring:
Full-text search
Captions are searchable once a BM25 index exists over them. Build it from (anchor, text) pairs:
query_text returns (anchor, score) tuples. The tokenizer is built in, so no model is involved — more on indexes.