DreamDB
DreamDB is a protocol for storing and searching multimodal data. Video, audio, text, and vectors share one timeline. Nothing is overwritten, vectors are indexed by where they sit in the address space, and reads come back as streams. It runs on any S3-compatible object store.
DreamDB has not reached 1.0. Read each capability's normative format and implementation limits separately; a Draft document can contain implemented sections. Backend credentials and policy remain deployment concerns, and the encrypted Object format has no reference Dataset implementation. See What's not there yet.
Python 0.0.10 and JavaScript 0.5.2 remain the published SDK versions. The Main / Unreleased guide covers newly merged entity keys, progressive geometry, ragged/CSR arrays and index capabilities. The 28-document specification index describes current contracts, not a promise that every feature is present in those registry packages.
Start building
- Install the SDKs
Python ingests, indexes, and trains. JavaScript reads and searches, in Node or the browser, and writes too.
- Write a dataset
Train a spatial index, declare a schema, append records. A local directory is a valid backend, so nothing else needs to be running.
- Query it from your app
Open the dataset over HTTP and run a vector, text, or scalar query. No server in between — the client fetches byte ranges and searches locally.
- Move it to real storage
Change the backend string to MinIO, S3, R2, or B2. Nothing else changes.
Four first principles
The rest of the spec follows from these four.
Time supports temporal lookup, not entity uniqueness. Main adds optional typed entity keys under a separate transaction contract.
Append-only and content-addressed. New data is layered on top, never written over.
Vector features encode into storage paths. A search is a coordinate calculation, not a traversal.
Encapsulation is streaming-native. The index is the player's seek pointer.
What you get
Nearest-neighbour search without a scan: the index is part of the address space.
Snapshot a dataset, branch it, merge work from several writers, reopen any earlier state.
Media is stored so it can be streamed, and the index doubles as a seek pointer into it.
Image, video, audio, embedding, and scalar fields in one dataset, on one timeline.
Any store with S3-style GET and PUT can hold a dataset. No server in between.
Tombstones hide records on read. Retained history remains readable; logical deletion is not selective erasure or a compliance guarantee.
SDKs and tools
dreamdb on PyPI. Ingest, index, snapshot, and delete, with a PyTorch IterableDataset and Arrow batches.
@dreamlake/dreamdb. The same Rust core compiled to WebAssembly: read, search, and write from Node or the browser.
Local disk, MinIO, S3, R2 — and what a browser needs to read them.
Install these docs into Claude, or point an agent at /llms.txt.
@dreamlake/dreamdb on npm and dreamdb on PyPI. Apache-2.0 / MIT.
Read the specification
The protocol is written down in full, and the SDKs are one implementation of it. If you want to know how something works underneath, or you're building your own client, start with the Spec overview and read forward from there: Data Model, Content Addressing, Time Encoding, Spatial Indexing.
What's not there yet
DreamDB is pre-1.0. What works is listed above and is in daily use, but a few things you might expect are not built:
- No access control. There are no users, roles, or capability tokens. Access is whatever your object store grants.
- No metrics or health endpoints. The SDKs log, but nothing is standardized for monitoring.
- Deletion is logical in the reference implementation. Tombstones hide records. The protocol defines physical reclamation, but the reference operator does not yet ship the
compact-tombstonespass; retained historical Manifests remain readable until retention and GC remove them. - No reference data-plane encryption yet. Spec 0019 defines the encrypted Object envelope and conformance rules, but the current SDKs do not write or read Lineage-v3 encrypted Objects. Backend-managed encryption is separate.
- Tested at about a million records. Larger benchmarks are still running, so treat bigger datasets as unproven.
Versions lists which SDK release to use and what stays compatible; release notes track the protocol itself.