API Reference
Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.
Selected public signatures are matched to the 0.0.12 release source. This page covers common pipeline operations; the feature guide adds query controls, entity-key, geometry and structured-array examples rather than claiming an exhaustive API inventory.
Looking for SQL? See SqlSession in the SQL guide. It is an unreleased source adapter, not part of the wheel version documented on this API page; the guide includes the required revision and local build steps.
Schema
Chainable field declarations. Every method returns the schema.
algorithm defaults to dreamdb.lsh-cosine, which requires a published spatial_index hash like every other partitioning algorithm — see Indexes. compressor takes a hash from publish_rabitq_compressor or train_and_publish_pq_compressor; rerank=True keeps full-precision vectors alongside compressed codes for a second scoring pass.
add_array declares a fixed-shape numeric value. Python accepts and returns NumPy arrays, validating the persisted dtype, shape, byte order, layout, codec, units, semantic type, frame, and axis order. The supported v1 dtypes are f32, f64, i8/i16/i32/i64, and u8/u16/u32/u64.
Dataset
Lifecycle
open recovers the schema from the manifest; pass one only if you want it checked field-for-field. open_at takes the dict snapshot() returned.
Writing
samples is a list of dicts keyed by field name. hot=True routes writes through a hot shard; threshold/age checks run during ingest, not in an independent background timer. Use flush_hot() when an explicit flush is needed. embedding_specs supplies per-field embedding identities, while vector queries use spec_id. ingest_video fragments a file with ffmpeg, so ffmpeg has to be on PATH.
Dataset.add_array adds an optional typed-array field to an existing dataset. add_constant supports the SDK's defined text, URI, SPDX, and JSON constant modalities; add_array_constant stores one typed NumPy array for the complete timeline.
Reading and querying
The iter_* methods yield dicts of columns plus _time_anchors. query_* methods return anchors, or (anchor, score) tuples for the scored ones. nprobe widens a vector search across neighbouring cells.
rerank_pool_size is a positive integer at least top_k; omitted, it preserves
the five-times-top-k partition pool. rerank_mode is inherit, approximate
or exact. Explicit modes override the Schema for this query only; exact mode
refuses missing exact capability rather than falling back. These controls do
not widen probes or guarantee global recall. Graph queries reject explicit pool
size or non-inherit mode. See query examples and limits.
The VideoItem metadata methods do not download media. item_key is opaque bytes and anchors remain exact Python integers. read_video_item_range returns initialization bytes and complete fragments overlapping an item-relative half-open range; it materializes the result in memory and never crosses into the next item.
The raw iter_stream fast path still requires an embedding field. A projection without one raises streaming iter with no embedding fields — use iter() for scalar-only:
Scalar-only datasets are otherwise usable: query_scalar, iter_scalar, distinct_values, list_anchors, count, and history all work without an embedding.
count() now works with every schema and excludes tombstoned anchors. It counts the union of distinct visible anchors across Tracks, so it is O(U_total) and materializes that anchor set rather than reading a stored counter.
iter_arrow_batches uses the embedding-only stream when exactly one embedding field is requested, and a windowed eager join for general projections. window_batches=None restores the fully eager behavior. Typed-array columns use Arrow FixedShapeTensor. The method needs pyarrow and NumPy, and says so when either is missing:
Versioning
snapshot returns {"label": "<ref>@<label>", "manifest": "<hash>", "timeline": "<hash>"} — keep it to reopen that exact state. Details in Versioning.
Indexes and layers
A layer attaches a derived field to an existing one — a second embedding of the same images, a thumbnail, a text index. compact returns counters:
Metadata
Module functions
Jobs that happen outside a dataset's lifetime.
Each train_and_publish_* returns a 33-byte content multihash to pass as spatial_index= or compressor=. They are deterministic: the same sample, parameters, and seed produce the same hash, so two workers publishing independently is a content-addressed no-op rather than a conflict.
gc defaults to dry_run=True. Destructive GC additionally requires
writers_paused=True after externally stopping and draining writers, uploads,
staged transactions and root changes, with a settled Ref listing. Keep that
fence for the whole run; the flag acquires no lock. Use roots (base32 Manifest
hashes) to retain otherwise unrooted reader snapshots. An empty root set refuses.
A concurrent dry-run is advisory, not a reusable deletion plan; age retention
cannot protect a concurrent writer reusing old Objects. Version 0.0.12 fixes
reference-mode vector-pool retention, but does not recover already-deleted data.
PyTorch
See Training pipelines.