API Reference
Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.
Signatures are from the package's own index.d.ts. Return shapes are what the calls actually produced when run against a live dataset — the wasm bindings type most payloads as any, so the types alone do not tell you enough.
Looking for SQL? See SqlSession in the SQL guide. It is an unreleased source adapter, not part of the package version documented on this API page; the guide covers Node, lazy browser loading and exact BigInt values.
Space
The read side. Open one with fromUri; there is no public constructor.
uri is a …/refs/<name> or …/manifests/<hash> URL. Pass null for backend to use direct fetch against the URI's base — enough for any readable bucket. Pass a backend object when reads need credentials or a proxy.
Properties
| Property | Type | Value |
|---|---|---|
manifestHash | string | The manifest this Space resolved to |
connectorBase | string | Storage root the URI was parsed against, e.g. https://bucket.example |
tracks / trackFor
coverage is a pair of bigint anchors. JSON.stringify throws on it — give it a replacer if you are logging tracks.
queryVector
opts accepts { topK?, probeCount?, specId?, rerankPoolSize?, rerankMode? },
or a bare number meaning topK. specId is the declared embedding identity in
base32. rerankMode is "inherit" (Schema default), "approximate" (no exact
rerank reads), or "exact" (refuse missing exact capability). rerankPoolSize
must be a positive u32 integer at least topK; omitted, the partition pool
defaults to five times top-k. Do not pass null as an omitted pool option.
Graph queries reject explicit pool size or non-inherit mode. The pool is not
probe coverage and exact scoring is not global exact recall. Returns hits
sorted by descending score:
anchor comes back as a number here, not a bigint. Raising probeCount widens the search across neighbouring buckets — worth doing on small datasets, where a query vector can land in a bucket that happens to be empty.
queryScalar
Every anchor whose field satisfies op value, tombstone-suppressed, ascending, deduped:
readScalarColumn
A whole scalar track as a Map, keyed by the anchor as a string:
Typed arrays
DecodedArray carries a JavaScript TypedArray in values plus the authoritative dtype, shape, endianness, layout, codec, semantic type, units, frame, and axis order. Per-item columns use decimal anchor strings as keys, matching readScalarColumn:
VideoItems
The two lookup methods return metadata without downloading media. readVideoItemRange returns the verified initialization segment and complete CMAF fragments overlapping an item-relative half-open interval. It materializes those selected bytes in memory; it is not a streaming playback API. All returned times are exact bigint nanoseconds.
Metadata and semantic cache
description() reads the manifest-pinned dataset description. The semantic cache defaults to 1 GiB; setting its capacity to 0 disables it, and the setter returns the number of bytes evicted.
resolveObjectIndex
The stored items of a track, as [startAnchor, endAnchor, size, hash] tuples — bigint, bigint, number, and a 33-byte multihash. This is how you locate blob payloads; see Browser usage for turning one into a URL.
queryText / queryHybrid
queryText runs BM25 over a text field, tokenizing with the built-in dreamdb.utf8-words tokenizer — no external model, even when the field's embeddings came from one. queryHybrid fuses a lexical and a dense sub-query per 0015; opts carries textField, textQuery, vectorField, topK, and fusion ('rrf' | 'linear' | 'max' | 'pareto'). Both return the same { anchor, score } shape as queryVector.
These two are documented from the type definitions and the spec. Unlike everything else on this page, they were not exercised against a live dataset while writing it — building one needs a text index, which the JavaScript SDK cannot create.
history
Walks the manifest DAG back from HEAD:
Any manifestHash from this list can be handed back to Space.fromUri as a …/manifests/<hash> URL, which is what time travel is.
Writer
Appends to a dataset that already exists. Present in every build, including the browser.
Samples are objects with an anchor and one entry per field, each naming its kind:
appendMany returns the count written. Nothing is visible to a reader until commit(), which advances the ref under compare-and-swap and returns the new manifest hash. snapshot(label) pins the current manifest under <ref>@<label>.
ingestCmaf is Node-only. It accepts an initialization segment plus already-fragmented [bytes, tStartNs, tEndNs] tuples, appends them to a video field, and publishes the new Track and manifest. It does not run ffmpeg or fragment a source file, and it holds the supplied clip bytes in memory. Repeated calls append clips only when their initialization segment is unchanged.
Authoring
Creates datasets and reshapes them. Node only — the browser build exports a stub whose methods throw.
Fields are declared as plain objects — there is no schema builder class:
Embedding fields declaring dreamdb.ivf-cosine or dreamdb.imi-cosine cannot be created here — those index families need training data that does not exist yet. Create with dreamdb.lsh-cosine and attach a trained index as a layer, or build the dataset in Python.
Typed-array fields use { kind: "array", dtype, shape, endianness?, layout?, codec?, semanticType?, units?, frame?, axisOrder?, trackKind?, objectKind? }. Event/unbucketed is the default. A constant declares trackKind: "constant" and objectKind: "constant", then stores its payload with addArrayConstant.
listRefs() needs a backend that implements list; without one it fails with connector: LIST refs/: operation not supported: backend.list.
Backends
A backend is any object with these methods. Only get is required, and a get-only backend is a read-only backend:
Write entry points check for put up front and refuse by name, rather than failing partway through a commit.
range is half-open — [start, end), the shape the Rust core uses. An HTTP Range header is inclusive, so a fetch-based backend must send bytes=${start}-${end - 1}. Sending end fetches one byte too many.
getStream is optional and does not change results. Implement it to keep full-object reads bounded when DreamDB must scan a historical oversized inline Track. Its chunks must themselves be bounded; one object-sized chunk is functionally valid but does not provide the bounded-memory profile.
S3Backend
Unsigned reads against a base URL. list is a stub that returns empty. Read-only: it has no put, so it cannot back a Writer or Authoring.
PresignedBackend
For writing from a browser without shipping credentials: your server mints short-lived presigned PUT URLs and performs the ref's compare-and-swap. See Browser usage.
Helpers
Pure functions over the protocol's encodings. All were run against live data except where noted.
| Function | Returns |
|---|---|
version() | Package version — '0.5.4' |
bytesToBase32(bytes) | Base32 address for a hash from resolveObjectIndex |
multihashBase32(bytes) / multihashHex(bytes) | Multihash renderings |
zeroSpatialKey() / ZERO_SPATIAL_KEY | '0000000000000000' — the spatial-key segment for non-bucketed objects |
timeAnchorHex(anchor) / timeAnchorFromHex(hex) | 16-char hex time anchors. Takes a bigint |
timeBucket(...) | Time-bucket arithmetic |
modalityParse(str) | { class, encoding, trackKind, objectKind, params } |
encodeCbor(value) / decodeCbor(bytes) | Canonical CBOR |
typedArrayDecode(itemType, payload) | A typed-array payload decoded to the same shape as readArrayColumn |
rangeHeader(...), addressRoundTrip(...), addressVariant(...), spatialKeyRoundTrip(...) | Address and range helpers used by the conformance suite |
Note that bytesToBase32 is what turns a resolveObjectIndex hash into an object address; multihashBase32 renders the same bytes differently and will 404 if used as a path segment.