DreamDB

Python SDK

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

dreamdb is where data gets in. It creates datasets, trains spatial indexes, ingests video, and feeds PyTorch — the parts of DreamDB the browser SDK deliberately doesn't carry.

bash
pip install dreamdb numpy

It is a compiled extension around the same Rust core as the JavaScript SDK, so the two write identical bytes. numpy is optional but every example uses it.

Two classes

Module-level functions cover the jobs that happen outside a dataset's lifetime: training indexes and compressors, garbage collection, and comparing refs.

What only Python can do

PythonJavaScript
Create a datasetAnywhereNode only
Train a spatial indexYesNo
dreamdb.ivf-cosine / imi-cosine fieldsYesCannot be created
Build a text (BM25) indexYesNo
Video ingest with fragmentingYes, via ffmpegNo
Typed numeric array fields and constantsYesYes, from 0.5.2
Read VideoItems by stable key or timeYesYes, from 0.5.2
PyTorch / Arrow batchesYesNo

If a pipeline writes data or builds an index, it wants this SDK. If an application reads and searches, either will do.

A local directory is a real backend

Nothing needs to be running to use DreamDB from Python:

python
ds = db.Dataset.create("photos", schema, backend="file:///tmp/photos")

memory:// exists for tests, and http(s):// for any S3-compatible endpoint. See Storage backends.

Where to go next