DreamDB

API Reference

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

Signatures are from the package's own index.d.ts. Return shapes are what the calls actually produced when run against a live dataset — the wasm bindings type most payloads as any, so the types alone do not tell you enough.

Looking for SQL? See SqlSession in the SQL guide. It is an unreleased source adapter, not part of the package version documented on this API page; the guide covers Node, lazy browser loading and exact BigInt values.

Space

The read side. Open one with fromUri; there is no public constructor.

ts
static fromUri(uri: string, backend: any): Promise<Space>

uri is a …/refs/<name> or …/manifests/<hash> URL. Pass null for backend to use direct fetch against the URI's base — enough for any readable bucket. Pass a backend object when reads need credentials or a proxy.

js
const space = await Space.fromUri("https://bucket.example/refs/photos", null);

Properties

PropertyTypeValue
manifestHashstringThe manifest this Space resolved to
connectorBasestringStorage root the URI was parsed against, e.g. https://bucket.example

tracks / trackFor

ts
tracks(): Array<ResolvedTrack>
trackFor(key_or_modality: string): ResolvedTrack
timelineId(): string
js
space.tracks()
// [ { modality: 'embedding.f32.dim=64.bucketed',
//     field: 'embedding', key: 'embedding',
//     kind: 'continuous', role: 'base',
//     address: 'dyen47624nsye67rgwtrjrqlt25dmam5us2caihtv4nbmcb3tvoos',
//     timeline: 'd33xiqqkreoyagxxtnwue2xjbb52okrit4nbul3mcxl4dptjyph5u',
//     coverage: [ 1785307993507390000n, 1785307993507689001n ] },
//   { modality: 'scalar.cat.field=label', field: 'label', key: 'label',
//     kind: 'event', role: 'base', … } ]

coverage is a pair of bigint anchors. JSON.stringify throws on it — give it a replacer if you are logging tracks.

queryVector

ts
queryVector(field: string, query: Float32Array, opts: any): Promise<Array<any>>

opts accepts { topK?, probeCount?, specId?, rerankPoolSize?, rerankMode? }, or a bare number meaning topK. specId is the declared embedding identity in base32. rerankMode is "inherit" (Schema default), "approximate" (no exact rerank reads), or "exact" (refuse missing exact capability). rerankPoolSize must be a positive u32 integer at least topK; omitted, the partition pool defaults to five times top-k. Do not pass null as an omitted pool option. Graph queries reject explicit pool size or non-inherit mode. The pool is not probe coverage and exact scoring is not global exact recall. Returns hits sorted by descending score:

js
await space.queryVector("embedding", vec, { topK: 3 })
// [ { anchor: 1785307993507656000, score: 0.3810671269893646 }, … ]

anchor comes back as a number here, not a bigint. Raising probeCount widens the search across neighbouring buckets — worth doing on small datasets, where a query vector can land in a bucket that happens to be empty.

queryScalar

ts
queryScalar(field: string, op: string, value: any): Promise<Array<any>>

Every anchor whose field satisfies op value, tombstone-suppressed, ascending, deduped:

js
await space.queryScalar("label", "eq", "cat")  // 100 anchors

readScalarColumn

ts
readScalarColumn(track: any, _opts: any): Promise<Map<any, any>>

A whole scalar track as a Map, keyed by the anchor as a string:

js
const labels = await space.readScalarColumn(space.trackFor("label"), null);
labels.size;                          // 300
labels.get("1785307993507390000");    // 'cat'

Typed arrays

ts
readArrayColumn(field: string): Promise<Map<string, DecodedArray>>
readArrayConstant(field: string): Promise<DecodedArray>

DecodedArray carries a JavaScript TypedArray in values plus the authoritative dtype, shape, endianness, layout, codec, semantic type, units, frame, and axis order. Per-item columns use decimal anchor strings as keys, matching readScalarColumn:

js
const poses = await space.readArrayColumn("pose");
const pose = poses.get("1735689600000000000");
// pose.values instanceof Float32Array; pose.shape is [2, 2]

const intrinsics = await space.readArrayConstant("intrinsics");

VideoItems

ts
videoItemAt(field: string, anchorNs: bigint | number): Promise<VideoItemInfo | undefined>
videoItemByKey(field: string, itemKey: Uint8Array): Promise<VideoItemInfo | undefined>
readVideoItemRange(
  field: string,
  itemKey: Uint8Array,
  relativeStartNs: bigint | number,
  relativeEndNs: bigint | number,
): Promise<VideoItemRead | undefined>

The two lookup methods return metadata without downloading media. readVideoItemRange returns the verified initialization segment and complete CMAF fragments overlapping an item-relative half-open interval. It materializes those selected bytes in memory; it is not a streaming playback API. All returned times are exact bigint nanoseconds.

Metadata and semantic cache

ts
description(): Promise<string | undefined>
semanticCacheCapacity(): Promise<number>
semanticCacheStats(): Promise<{ entries: number; bytesUsed: number; maxBytes: number }>
setSemanticCacheCapacity(nBytes: number): Promise<number>

description() reads the manifest-pinned dataset description. The semantic cache defaults to 1 GiB; setting its capacity to 0 disables it, and the setter returns the number of bytes evicted.

resolveObjectIndex

ts
resolveObjectIndex(track: any): Promise<Array<any>>

The stored items of a track, as [startAnchor, endAnchor, size, hash] tuples — bigint, bigint, number, and a 33-byte multihash. This is how you locate blob payloads; see Browser usage for turning one into a URL.

js
const [start, end, size, hash] = (await space.resolveObjectIndex(track))[0];
// 1785315427559340000n, 1785315427559340001n, 160, Uint8Array(33)

queryText / queryHybrid

ts
queryText(field: string, query: string, top_k: number): Promise<Array<any>>
queryHybrid(vector: Float32Array, opts: any): Promise<Array<any>>

queryText runs BM25 over a text field, tokenizing with the built-in dreamdb.utf8-words tokenizer — no external model, even when the field's embeddings came from one. queryHybrid fuses a lexical and a dense sub-query per 0015; opts carries textField, textQuery, vectorField, topK, and fusion ('rrf' | 'linear' | 'max' | 'pareto'). Both return the same { anchor, score } shape as queryVector.

These two are documented from the type definitions and the spec. Unlike everything else on this page, they were not exercised against a live dataset while writing it — building one needs a text index, which the JavaScript SDK cannot create.

history

ts
history(max_depth?: number | null): Promise<Array<any>>

Walks the manifest DAG back from HEAD:

js
(await space.history(3))[0]
// { manifestHash: 'd2wn2nmxvqlmt7j6cvvd3r2xml22bv2errxeahfqmoshdjjjzdfxe',
//   ts: 1785307993513719000,
//   writer: 'dreamdb-dataset/0.1#photos',
//   tracks: [ { modality: 'embedding.f32.dim=64.bucketed', address: '…' }, … ] }

Any manifestHash from this list can be handed back to Space.fromUri as a …/manifests/<hash> URL, which is what time travel is.

Writer

Appends to a dataset that already exists. Present in every build, including the browser.

ts
static open(uri: string, backend: any): Promise<Writer>
appendMany(samples: Array<any>): Promise<number>
commit(): Promise<string>
snapshot(label: string): Promise<string>
deleteRecords(anchors: BigUint64Array, reason?: string | null): Promise<string>
ingestCmaf(field: string, init: Uint8Array, fragments: CmafFragment[]): Promise<CmafIngestResult>
readonly refName: string
readonly manifestHash: string

Samples are objects with an anchor and one entry per field, each naming its kind:

js
await writer.appendMany([{
  anchor: BigInt(Date.now()) * 1_000_000n,
  vec: { kind: "embedding", algorithm: "dreamdb.lsh-cosine", vector: new Float32Array(16) },
  tag: { kind: "categorical", value: "a" },
}]);

appendMany returns the count written. Nothing is visible to a reader until commit(), which advances the ref under compare-and-swap and returns the new manifest hash. snapshot(label) pins the current manifest under <ref>@<label>.

ingestCmaf is Node-only. It accepts an initialization segment plus already-fragmented [bytes, tStartNs, tEndNs] tuples, appends them to a video field, and publishes the new Track and manifest. It does not run ffmpeg or fragment a source file, and it holds the supplied clip bytes in memory. Repeated calls append clips only when their initialization segment is unchanged.

Authoring

Creates datasets and reshapes them. Node only — the browser build exports a stub whose methods throw.

ts
static create(name: string, schema: Array<any>, base: string, backend: any): Promise<Authoring>
static open(uri: string, backend: any): Promise<Authoring>
writer(): Writer
branch(new_name: string): Promise<Authoring>
mergeMany(branches: string[]): Promise<string>
compact(modality: string | null | undefined, threshold: number, max_cells: number): Promise<any>
addField(field: any): Promise<void>
addArrayConstant(field: any, payload: Uint8Array): Promise<void>
addEmbeddingLayer(name, parent_field, dim, algorithm, spatial_index, compressor, …): Promise<void>
addImageLayer(name: string, parent_field: string, samples: Array<any>, mime: string): Promise<void>
addScalarLayer(name: string, parent_field: string, value_type: string, samples: Array<any>): Promise<void>
listRefs(): Promise<string[]>
deleteRef(name: string): Promise<void>
setDescription(text: string): Promise<void>
readonly refName: string

Fields are declared as plain objects — there is no schema builder class:

js
await Authoring.create("js-notes", [
  { name: "vec", kind: "embedding", dim: 16, algorithm: "dreamdb.lsh-cosine" },
  { name: "tag", kind: "scalar", valueType: "categorical", required: false },
], base, backend);

Embedding fields declaring dreamdb.ivf-cosine or dreamdb.imi-cosine cannot be created here — those index families need training data that does not exist yet. Create with dreamdb.lsh-cosine and attach a trained index as a layer, or build the dataset in Python.

Typed-array fields use { kind: "array", dtype, shape, endianness?, layout?, codec?, semanticType?, units?, frame?, axisOrder?, trackKind?, objectKind? }. Event/unbucketed is the default. A constant declares trackKind: "constant" and objectKind: "constant", then stores its payload with addArrayConstant.

listRefs() needs a backend that implements list; without one it fails with connector: LIST refs/: operation not supported: backend.list.

Backends

A backend is any object with these methods. Only get is required, and a get-only backend is a read-only backend:

ts
interface Backend {
  get(path: string, range?: { start: number; end: number })
    : Promise<Uint8Array | { bytes: Uint8Array; etag?: string; totalLength?: number }>
  getStream?(path: string)
    : Promise<ReadableStream<Uint8Array> | {
        stream: ReadableStream<Uint8Array>;
        etag?: string;
        totalLength?: number;
      }>
  head?(path: string): Promise<{ exists?: boolean; etag?: string; size?: number }>
  put?(path: string, bytes: Uint8Array,
       opts?: { ifMatch?: string; ifNoneMatchStar?: boolean })
    : Promise<{ status: "created" | "exists" | "casFailed"; etag?: string }>
  delete?(path: string): Promise<void>
  list?(prefix: string): Promise<string[]>
}

Write entry points check for put up front and refuse by name, rather than failing partway through a commit.

range is half-open — [start, end), the shape the Rust core uses. An HTTP Range header is inclusive, so a fetch-based backend must send bytes=${start}-${end - 1}. Sending end fetches one byte too many.

getStream is optional and does not change results. Implement it to keep full-object reads bounded when DreamDB must scan a historical oversized inline Track. Its chunks must themselves be bounded; one object-sized chunk is functionally valid but does not provide the bounded-memory profile.

S3Backend

ts
constructor(base_url: string)
get(path: string, _opts: any): Promise<Uint8Array>
getStream(path: string): Promise<{ stream: ReadableStream<Uint8Array>; totalLength?: number }>
list(_prefix: string): Promise<Array<any>>

Unsigned reads against a base URL. list is a stub that returns empty. Read-only: it has no put, so it cannot back a Writer or Authoring.

PresignedBackend

ts
constructor(options: PresignedBackendOptions)
get(path, range?): Promise<{ bytes: Uint8Array; etag?: string }>
head(path): Promise<{ exists?: boolean; etag?: string; size?: number }>
put(path, bytes, opts?): Promise<{ status: PutStatus; etag?: string }>

For writing from a browser without shipping credentials: your server mints short-lived presigned PUT URLs and performs the ref's compare-and-swap. See Browser usage.

Helpers

Pure functions over the protocol's encodings. All were run against live data except where noted.

FunctionReturns
version()Package version — '0.5.4'
bytesToBase32(bytes)Base32 address for a hash from resolveObjectIndex
multihashBase32(bytes) / multihashHex(bytes)Multihash renderings
zeroSpatialKey() / ZERO_SPATIAL_KEY'0000000000000000' — the spatial-key segment for non-bucketed objects
timeAnchorHex(anchor) / timeAnchorFromHex(hex)16-char hex time anchors. Takes a bigint
timeBucket(...)Time-bucket arithmetic
modalityParse(str){ class, encoding, trackKind, objectKind, params }
encodeCbor(value) / decodeCbor(bytes)Canonical CBOR
typedArrayDecode(itemType, payload)A typed-array payload decoded to the same shape as readArrayColumn
rangeHeader(...), addressRoundTrip(...), addressVariant(...), spatialKeyRoundTrip(...)Address and range helpers used by the conformance suite
js
modalityParse("embedding.f32.dim=64.bucketed")
// { class: 'embedding', encoding: 'f32',
//   trackKind: 'Continuous', objectKind: 'SpatialBucket', params: { dim: '64' } }

Note that bytesToBase32 is what turns a resolveObjectIndex hash into an object address; multihashBase32 renders the same bytes differently and will 404 if used as a path segment.