DreamDB

Versioning

Every write produces a new immutable manifest pointing at the previous one. Snapshots, branches, and time travel are all just names for reading a particular manifest — none of them copies data.

Snapshots

python
version = ds.snapshot("v1")
# {'label': 'notes@v1',
#  'manifest': 'dyysbk4jlsq5sycez7omujhwtuiydgdlrav4ushkekynxcgnqptli',
#  'timeline': 'd33xiqqkreoyagxxtnwue2xjbb52okrit4nbul3mcxl4dptjyph5u'}

Keep that dict — it is how you get back:

python
old = db.Dataset.open_at(version, backend=BACKEND)
old.count()      # the count as of v1, whatever has happened since

A snapshot is a label on a manifest, so it costs one small object and is instant regardless of dataset size. Reproducible training runs pin one.

Branches

python
worker = ds.branch("worker-0")
worker.append_many(rows)

A branch starts from the current manifest and advances independently. Both refs exist side by side:

python
ds.list_refs()      # ['notes', 'notes@v1', 'worker-0', 'worker-1']

Sharded ingest

The reason branches exist. Split a large ingest across processes or machines, each on its own branch, then merge:

python
# each worker, independently
b = ds.branch(f"worker-{i}")
b.append_many(my_slice)

# once they finish
ds.merge_many(["worker-0", "worker-1"])

Verified against 0.0.10 with a trunk of 10 records and two workers adding 5 each:

worker-0: 15
worker-1: 15
merge_many → d3vcqzcwfjqx…
trunk after merge: 20
at v1: 10

Workers never coordinate. They write to disjoint refs, and the merge is a manifest operation over content-addressed objects — no data moves.

Merging refuses when two branches bound the same modality to divergent index lineages, rather than mixing vectors quantized under incompatible codebooks:

config: refusing union-merge: embedding modality `text.utf8.bm25` is bound to
divergent index lineage on the two branches … Merging would mix vectors
indexed/quantized under incompatible SpatialIndexes or codebooks, silently
corrupting query results (spec/0008 §6.5 MUST-REFUSE).

In practice this is what stops a text index and sharded ingest from combining today — build the index after the merge.

History

python
for entry in ds.history(10):
    print(entry["manifest"], entry["ts"], entry["writer"])

Each entry names its parent, so the list is the ref's ancestry. Any manifest hash in it can be reopened with Dataset.open_by_manifest.

Deletion

python
ds.delete(anchors, reason="gdpr")
ds.count()          # drops by len(anchors)
ds.tombstone_set()  # the suppressed anchors

A tombstone suppresses records on read. History stays intact — that is what makes time travel and merges sound — so the bytes are still on disk until a future compaction step reclaims them. If your obligation is that the bytes are gone, this is not yet that; see Tombstones.

Comparing two refs

python
db.compare_refs(["notes", "worker-0"], fields=["label"], backend=BACKEND)

Useful before a merge, or to see what a branch actually changed.