DreamDB

DreamDB

DreamDB is a protocol for storing and searching multimodal data. Video, audio, text, and vectors share one timeline. Nothing is overwritten, vectors are indexed by where they sit in the address space, and reads come back as streams. It runs on any S3-compatible object store.

Pre-1.0

DreamDB has not reached 1.0. Read each capability's normative format and implementation limits separately; a Draft document can contain implemented sections. Backend credentials and policy remain deployment concerns, and the encrypted Object format has no reference Dataset implementation. See What's not there yet.

Main is ahead of the released packages

Python 0.0.10 and JavaScript 0.5.2 remain the published SDK versions. The Main / Unreleased guide covers newly merged entity keys, progressive geometry, ragged/CSR arrays and index capabilities. The 28-document specification index describes current contracts, not a promise that every feature is present in those registry packages.

Start building

  1. Install the SDKs

    Python ingests, indexes, and trains. JavaScript reads and searches, in Node or the browser, and writes too.

    bash
    pip install dreamdb numpy
    npm install @dreamlake/dreamdb

    Full installation guide

  2. Write a dataset

    Train a spatial index, declare a schema, append records. A local directory is a valid backend, so nothing else needs to be running.

    Quick Start

  3. Query it from your app

    Open the dataset over HTTP and run a vector, text, or scalar query. No server in between — the client fetches byte ranges and searches locally.

    JavaScript SDK · Read-only SQL

  4. Move it to real storage

    Change the backend string to MinIO, S3, R2, or B2. Nothing else changes.

    Storage backends


Four first principles

The rest of the spec follows from these four.

Time is the primary clustering axis

Time supports temporal lookup, not entity uniqueness. Main adds optional typed entity keys under a separate transaction contract.

Immutability is the bedrock of collaboration

Append-only and content-addressed. New data is layered on top, never written over.

Retrieval is localization, not scanning

Vector features encode into storage paths. A search is a coordinate calculation, not a traversal.

Data is stream

Encapsulation is streaming-native. The index is the player's seek pointer.


What you get


SDKs and tools


Read the specification

The protocol is written down in full, and the SDKs are one implementation of it. If you want to know how something works underneath, or you're building your own client, start with the Spec overview and read forward from there: Data Model, Content Addressing, Time Encoding, Spatial Indexing.

What's not there yet

DreamDB is pre-1.0. What works is listed above and is in daily use, but a few things you might expect are not built:

  • No access control. There are no users, roles, or capability tokens. Access is whatever your object store grants.
  • No metrics or health endpoints. The SDKs log, but nothing is standardized for monitoring.
  • Deletion is logical in the reference implementation. Tombstones hide records. The protocol defines physical reclamation, but the reference operator does not yet ship the compact-tombstones pass; retained historical Manifests remain readable until retention and GC remove them.
  • No reference data-plane encryption yet. Spec 0019 defines the encrypted Object envelope and conformance rules, but the current SDKs do not write or read Lineage-v3 encrypted Objects. Backend-managed encryption is separate.
  • Tested at about a million records. Larger benchmarks are still running, so treat bigger datasets as unproven.

Versions lists which SDK release to use and what stays compatible; release notes track the protocol itself.


Next