DreamDB

API Reference

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

Selected public signatures are matched to the 0.0.12 release source. This page covers common pipeline operations; the feature guide adds query controls, entity-key, geometry and structured-array examples rather than claiming an exhaustive API inventory.

Looking for SQL? See SqlSession in the SQL guide. It is an unreleased source adapter, not part of the wheel version documented on this API page; the guide includes the required revision and local build steps.

Schema

Chainable field declarations. Every method returns the schema.

python
Schema.add_image(name, mime="jpeg", required=True, chunk_size=None, pack_items=None)
Schema.add_video(name, mime="mp4", required=True, chunk_size=None, pack_items=None)
Schema.add_audio(name, mime="wav", required=True, chunk_size=None, pack_items=None)
Schema.add_embedding(name, dim, algorithm="dreamdb.lsh-cosine", required=True,
                     lsh_bits=None, compressor=None, spatial_index=None, rerank=False,
                     bucket_max_bytes=None, embed_spec=None)
Schema.add_geometry(name, family="point")
Schema.add_ragged_array(name, dtype, inner_shape=(), **kwargs)
Schema.add_sparse_csr(name, dtype, shape, **kwargs)
Schema.add_array(name, dtype, shape, *, endianness="little", layout="c",
                 codec="raw", semantic_type="none", units="none", frame="none",
                 axis_order="none", track_kind="event", object_kind="unbucketed",
                 required=True)
Schema.add_scalar_categorical(name, required=True)
Schema.add_scalar_string(name, required=True)
Schema.add_scalar_int(name, required=True)
Schema.add_scalar_float(name, required=True)
Schema.add_scalar_bool(name, required=True)
Schema.add_scalar_timestamp(name, required=True)

algorithm defaults to dreamdb.lsh-cosine, which requires a published spatial_index hash like every other partitioning algorithm — see Indexes. compressor takes a hash from publish_rabitq_compressor or train_and_publish_pq_compressor; rerank=True keeps full-precision vectors alongside compressed codes for a second scoring pass.

add_array declares a fixed-shape numeric value. Python accepts and returns NumPy arrays, validating the persisted dtype, shape, byte order, layout, codec, units, semantic type, frame, and axis order. The supported v1 dtypes are f32, f64, i8/i16/i32/i64, and u8/u16/u32/u64.

Dataset

Lifecycle

python
Dataset.create(name: str, schema, backend: str) -> Dataset
Dataset.open(name: str, schema=None, backend: str = "") -> Dataset
Dataset.open_at(version: dict, backend: str) -> Dataset
Dataset.open_ref(ref_name: str, backend: str) -> Dataset
Dataset.open_by_manifest(manifest_hash: str, backend: str) -> Dataset

open recovers the schema from the manifest; pass one only if you want it checked field-for-field. open_at takes the dict snapshot() returned.

Writing

python
append_many(samples, commit: bool = True, hot: bool = False, embedding_specs=None) -> int
add_array(name, dtype, shape, *, ..., required=False) -> Dataset
add_constant(field_name: str, modality: str, value) -> None
get_constant(field_name: str)
add_array_constant(field_name: str, value, *, dtype, shape, ...) -> None
get_array_constant(field_name: str)
commit() -> None
delete(anchors, reason: str | None = None) -> str
configure_hot_shard(flush_threshold: int = 10_000, ttl_seconds: int = 30) -> None
flush_hot() -> None
ingest_video(field: str, path: str, *, frag_duration: float = 2.0,
             anchor: int = 0, preview: dict | None = None) -> dict

samples is a list of dicts keyed by field name. hot=True routes writes through a hot shard; threshold/age checks run during ingest, not in an independent background timer. Use flush_hot() when an explicit flush is needed. embedding_specs supplies per-field embedding identities, while vector queries use spec_id. ingest_video fragments a file with ffmpeg, so ffmpeg has to be on PATH.

Dataset.add_array adds an optional typed-array field to an existing dataset. add_constant supports the SDK's defined text, URI, SPDX, and JSON constant modalities; add_array_constant stores one typed NumPy array for the complete timeline.

Reading and querying

python
iter_vector(field: str, query, top_k: int, batch_size: int = 64,
            where_eq=None, where=None, as_numpy: bool = False,
            nprobe: int | None = None, fields=None, spec_id=None,
            rerank_pool_size: int | None = None, rerank_mode: str = "inherit")
iter_scalar(where_eq=None, where=None, batch_size: int = 64)
iter_stream(batch_size: int = 256, fields=None, channel_buffer: int = 8,
            start_ns: int = 0, end_ns: int | None = None)
iter_all_batches(...)
iter_arrow_batches(batch_size: int = 256, fields=None,
                   shuffle_seed: int | None = None, window_batches: int | None = 256)
query_scalar(field: str, op: str, value) -> list[int]
query_text(field: str, query: str, top_k: int = 10) -> list[tuple[int, float]]
query_hybrid(...)
query_phrase(...)
get_by_anchor(field: str, root: bytes, anchor: int)
video_item_by_key(field: str, item_key) -> dict | None
video_item_at(field: str, anchor: int) -> dict | None
read_video_item_range(field: str, item_key,
                      relative_start: int, relative_end: int) -> dict | None
distinct_values(field: str) -> list[tuple]
count() -> int
list_anchors() -> list[int]

The iter_* methods yield dicts of columns plus _time_anchors. query_* methods return anchors, or (anchor, score) tuples for the scored ones. nprobe widens a vector search across neighbouring cells.

rerank_pool_size is a positive integer at least top_k; omitted, it preserves the five-times-top-k partition pool. rerank_mode is inherit, approximate or exact. Explicit modes override the Schema for this query only; exact mode refuses missing exact capability rather than falling back. These controls do not widen probes or guarantee global recall. Graph queries reject explicit pool size or non-inherit mode. See query examples and limits.

The VideoItem metadata methods do not download media. item_key is opaque bytes and anchors remain exact Python integers. read_video_item_range returns initialization bytes and complete fragments overlapping an item-relative half-open range; it materializes the result in memory and never crosses into the next item.

The raw iter_stream fast path still requires an embedding field. A projection without one raises streaming iter with no embedding fields — use iter() for scalar-only:

python
ds.iter_stream(fields=["emb", "frame_idx"])   # works
ds.iter_stream(fields=["emb"])                # works
ds.iter_stream(fields=["frame_idx", "split"]) # raises, however many fields

Scalar-only datasets are otherwise usable: query_scalar, iter_scalar, distinct_values, list_anchors, count, and history all work without an embedding.

count() now works with every schema and excludes tombstoned anchors. It counts the union of distinct visible anchors across Tracks, so it is O(U_total) and materializes that anchor set rather than reading a stored counter.

iter_arrow_batches uses the embedding-only stream when exactly one embedding field is requested, and a windowed eager join for general projections. window_batches=None restores the fully eager behavior. Typed-array columns use Arrow FixedShapeTensor. The method needs pyarrow and NumPy, and says so when either is missing:

iter_arrow_batches requires `pyarrow`. Install with `pip install pyarrow`
or use `iter_vector` for the plain-dict path.

Versioning

python
snapshot(label: str) -> dict
branch(new_name: str) -> Dataset
merge(other_ref_name: str, strategy: str = "fast-forward") -> str
merge_many(branches) -> str
history(max_depth: int = 50)
current_manifest()
timeline()
ref_name()
list_refs()
delete_ref(name: str) -> None
tombstone_set() -> list[int]

snapshot returns {"label": "<ref>@<label>", "manifest": "<hash>", "timeline": "<hash>"} — keep it to reopen that exact state. Details in Versioning.

Indexes and layers

python
add_text_index(field: str, parent_field: str, docs: list[tuple[int, str]],
               k1: float = 1.5, b: float = 0.75, delta: float = 1.0, algorithm: str = ...)
add_embedding_layer(...)
add_image_layer(...)
append_embedding_layer(...)
build_graph_index(...)
build_anchor_index(field: str) -> bytes
add_multi_vector_index(...)
add_splade_index(...)
compact(modality: str | None = None, threshold: int = 1, max_cells: int = 0) -> dict

A layer attaches a derived field to an existing one — a second embedding of the same images, a thumbnail, a text index. compact returns counters:

python
ds.compact()
# {'cells_examined': 21, 'cells_compacted': 0, 'fragments_collapsed': 0, …}

Metadata

python
set_meta(key: str, value: str) -> None
meta() -> dict
set_description(text: str) -> None
description() -> str | None
set_coverage_complete(...)

Module functions

Jobs that happen outside a dataset's lifetime.

python
train_and_publish_ivf_centroids(backend, sample, dim, k, iterations=10, seed=None) -> bytes
train_and_publish_imi_centroids(backend, sample, dim, k_sub, iterations=10, seed=None) -> bytes
train_and_publish_pq_compressor(backend, sample, dim, m, k, iterations=10, seed=None) -> bytes
publish_rabitq_compressor(backend, dim, seed=None, bits_per_dim=1,
                          with_correction_factors=True) -> bytes
train_and_publish_vamana_graph(backend, sample, dim, modality, r=32, alpha=1.2,
                               l_build=100, ...) -> bytes
gc(backend, keep_manifests=10, keep_since_secs=0, dry_run=True,
   *, roots=None, writers_paused=False) -> None
compare_refs(refs: list[str], fields: list[str], backend: str, batch_size=256)

Each train_and_publish_* returns a 33-byte content multihash to pass as spatial_index= or compressor=. They are deterministic: the same sample, parameters, and seed produce the same hash, so two workers publishing independently is a content-addressed no-op rather than a conflict.

gc defaults to dry_run=True. Destructive GC additionally requires writers_paused=True after externally stopping and draining writers, uploads, staged transactions and root changes, with a settled Ref listing. Keep that fence for the whole run; the flag acquires no lock. Use roots (base32 Manifest hashes) to retain otherwise unrooted reader snapshots. An empty root set refuses. A concurrent dry-run is advisory, not a reusable deletion plan; age retention cannot protect a concurrent writer reusing old Objects. Version 0.0.12 fixes reference-mode vector-pool retention, but does not recover already-deleted data.

PyTorch

python
from dreamdb.torch import DreamDBIterableDataset

DreamDBIterableDataset(dataset, *, field, query, top_k, batch_size=64,
                       where_eq=None, as_numpy=True, transform=None)

See Training pipelines.