DreamDB

Quick Start

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

Write 300 records with Python, then query them from JavaScript. Everything runs on your machine, against a plain directory — no S3 account, no Docker, about ten minutes.

You need both SDKs installed: pip install dreamdb numpy and npm install @dreamlake/dreamdb.

1. Write a dataset

DreamDB stores vectors by where they land in space, so a spatial index has to exist before the first record does. Training it is one call, and the whole script is one file:

python
import numpy as np
import dreamdb as db

BACKEND = "file:///tmp/dreamdb-quickstart"
rng = np.random.default_rng(0)

# 1. Train a spatial index. It decides where each vector lands in storage.
sample = rng.standard_normal((512, 64)).astype(np.float32)
index = db.train_and_publish_ivf_centroids(BACKEND, sample, dim=64, k=16)

# 2. Declare the fields.
schema = (db.Schema()
          .add_embedding("embedding", dim=64,
                         algorithm="dreamdb.ivf-cosine", spatial_index=index)
          .add_scalar_categorical("label"))

# 3. Create the dataset and append 300 records.
ds = db.Dataset.create("photos", schema, backend=BACKEND)
ds.append_many([
    {"embedding": rng.standard_normal(64).astype(np.float32),
     "label": ["cat", "dog", "bird"][i % 3]}
    for i in range(300)
])

print("records:", ds.count())
print("snapshot:", ds.snapshot("v1")["label"])

# 4. Query it.
query = rng.standard_normal(64).astype(np.float32)
for batch in ds.iter_vector(field="embedding", query=query, top_k=3, batch_size=3):
    print("nearest:", batch["label"])

Create the directory and run it:

bash
mkdir -p /tmp/dreamdb-quickstart
python ingest.py
records: 300
snapshot: photos@v1
nearest: ['cat', 'dog', 'cat']

That directory is now a complete dataset: a ref, a manifest chain, one track per field, and the spatial index. Nothing else is involved — no server, no database process.

Real embeddings work the same way. Swap the random vectors for CLIP, or any model's output, and set dim to match. A 512-dimension CLIP field is add_embedding("clip", dim=512, ...).

2. Serve the directory

The JavaScript SDK reads over HTTP, not off disk, so point any static server at the directory. Browsers additionally need CORS headers, which --cors sends:

bash
npx http-server /tmp/dreamdb-quickstart -p 8791 --cors

In production this is S3, MinIO, or R2 instead — same protocol, same code. See Storage backends.

Serve it with something that honours Range. DreamDB reads vectors at byte offsets. A server that ignores Range and returns the whole object still gives correct results — the connectors slice locally — but it downloads the entire object on every ranged read, which stops being viable as soon as the data is large. python3 -m http.server ignores Range; http-server honours it.

3. Read it from JavaScript

js
import { Space } from "@dreamlake/dreamdb";

const space = await Space.fromUri("http://localhost:8791/refs/photos", undefined);

console.log("manifest:", space.manifestHash.slice(0, 12) + "…");
console.log("tracks:", space.tracks().map(t => t.key).join(", "));

const query = new Float32Array(64).map(() => Math.random() - 0.5);
const hits = await space.queryVector("embedding", query, { topK: 3 });

for (const hit of hits) {
  console.log(`anchor ${hit.anchor}  score ${hit.score.toFixed(3)}`);
}
bash
node read.mjs
manifest: d2wn2nmxvqlm…
tracks: embedding, label
anchor 1785307993507438000  score 0.352
anchor 1785307993507404000  score 0.323
anchor 1785307993507480000  score 0.299

The same code runs in a browser. Nothing sits between the page and storage — the SDK fetches byte ranges and does the search locally.

The JavaScript SDK can also write, given a backend that accepts PUT. A static directory server is not one; see Storage backends.

What you just built

  • Anchors are the primary index axis. 1785243621970432000 is a nanosecond time anchor used for ordering, joins, and ordinary record lookup. Model an independent entity id as a scalar field; VideoItem additionally carries an opaque stable item key.
  • The manifest is the version. Every write produced a new immutable manifest pointing at the previous one. snapshot("v1") labelled one of them; that label will resolve to those exact bytes forever.
  • The query never scanned. The trained index turned your query vector into a storage location, and the SDK fetched only the buckets that could contain neighbours.

Open it again later

A dataset outlives the process that wrote it. Reopen it by name and the schema comes back from the manifest:

python
ds = db.Dataset.open("photos", backend="file:///tmp/dreamdb-quickstart")
print(ds.count())

A snapshot reopens the same way, at the state it labelled:

python
old = db.Dataset.open_at(version, backend="file:///tmp/dreamdb-quickstart")

Next steps