DreamDB

Feature guide — Python 0.0.12 / npm 0.5.4

This guide targets Python 0.0.12 and npm 0.5.4. The core baseline is a119479. The entity, geometry and array examples originated in the preceding release; the new query controls and GC repair are described below. Older packages do not necessarily expose these options. Rust-only operator capabilities are called out separately, not presented as Python or JavaScript methods.

For read-only SQL v1, see the SQL guide: merged Rust API plus unreleased Python, npm and CLI source adapters. These are separate from the SDK releases described here; a source implementation is not a package release.

Upgrade collectors before reference-mode GC

Version 0.0.12 includes the fix for GC deleting VectorStorage referenced only through reference-mode Bucket records (#365). Both inline and paged indexes now use authenticated dependency traversal; malformed metadata refuses before sweep. Upgrade or fence every collector processing these roots. Updating only a reader does not make an old collector safe. The repair cannot recover previously deleted Objects or repair #89's historical Manifest metadata damage.

GC still requires an externally enforced maintenance window and retained reader roots. An age threshold and a prior dry-run do not substitute for that fence. See GC requirements.

Query-local exactness and pool size

For an existing partition-vector field and a query of the declared dimension:

python
batches = ds.iter_vector(
    "embedding", query, top_k=10,
    rerank_pool_size=50, rerank_mode="exact",
)
js
const hits = await space.queryVector("embedding", query, {
  topK: 10, rerankPoolSize: 50, rerankMode: "exact",
});

These are call fragments; use the Python or JavaScript quickstart to create/open the dataset. inherit keeps the Schema policy; approximate skips the exact rerank pass; exact requires an exact source for every selected candidate, with no silent fallback. Pool size is a positive integer at least top-k (default five times top-k), selected after eligibility filtering. It does not widen probe coverage or guarantee global exact recall. Graph queries reject an explicit pool or non-inherit mode. Raw L2 vectors retain their exact metric scores in all modes.

Rust operator APIs: not new SDK wrappers

The source release also includes fragmentation observations, explicit maintenance plans, bounded Fragment repacking, resumable Reencode admission, pull-based Ref Watch, create-only snapshot tags, media Stream selection and verified closure-copy plans. These are Rust Dataset/operator APIs; no corresponding new Python or JavaScript wrappers are implied. See release notes and the owning operations, evolution and copy-plan contracts. None adds a resident maintenance scheduler or an automatic business rollout.

What changed and what did not

CapabilityMain behaviorImportant boundary
Entity keys (#322)Optional string/bytes/int keys, stable entity ID, revision-checked atomic create/update/deleteTimeline-scoped; v1 key index ≤16 MiB; divergent key-index merges refuse even for disjoint edits
Geometry (#316)Point/splat cells, packed data, paged hierarchy, independent mesh LODs, authenticated progressive readsOne outer Item, not one Sample per point; metadata/cache and producer memory are not universally bounded
Ragged and CSR (#323)Dataset/Python/WASM writers and typed reads, not just codec declarationsStructured Arrow projection unsupported; dense/ragged/sparse remain different declared families
L2 (#321)Magnitude-preserving Euclidean partition/search pathDo not normalize inputs as cosine vectors or interpret its scores as cosine similarity
Compressed graphs (#320)ADC traversal with exact-source pool rerankingExact reranking cannot recover nodes excluded from candidate selection
Federation routing (#319)Exact-bound shard router over explicitly eligible childrenDoes not discover backends or guarantee global exact recall
Fresh graphs (#318)Append plus id-preserving consolidation; changed pages published, others reusedRaw graph lineage only; explicit calls, no resident scheduler or physical erasure
HotShards (#324/#328)Field-qualified hot embedding, scalar and text dataFlush age checked by ingest is not an independent timer; configure before depending on hot behavior
Embedding identity (#317, PR #327)spec_id checked at append, query and index-build boundariesApproximate compatibility never authorizes reusing a vector from another identity
Hybrid planning (#326)Executed plan, configurable latency estimate/budget, observable traceCost estimates are not a hard wall-clock SLA; no silent unfiltered fallback
Scalar B-trees (#325)Persistent path-copy updatesDoes not mean all indexes or all maintenance operations are constant-memory

The same Rust implementation underlies the SDKs. Shared-code tests are not independent implementations agreeing on a protocol.

Entity keys: keep the revision token

Python 0.0.12:

python
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-keys-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create(
        "keys", db.Schema().add_scalar_string("text"), backend=backend
    )
    ds.enable_entity_keys("string")
    first = ds.create_entity("document-1", {"text": "first"})
    second = ds.supersede_entity(
        "document-1", {"text": "second"}, expected=first["revision"]
    )
    assert second["entity_id"] == first["entity_id"]
    reopened = db.Dataset.open("keys", backend=backend)
    assert reopened.get_entity("document-1")["sample"] == {"text": "second"}

A stale expected revision must refuse, not overwrite the newer value. Keep the returned token as bytes; do not replace it with a timestamp or decoded text. Logical deletion leaves historical snapshots readable. Deleted entries can be explicitly restored using their latest revision; an ordinary scalar filter does not provide any of these transaction guarantees. Read spec/0026 before enabling keyed writes.

For a WASM 0.5.4 writer, the corresponding calls are enableEntityKeys, createEntity, supersedeEntity, deleteEntity and upsertEntity. An existing Space is pinned: reopen to read a later publication. Given an initialized Writer named writer (from Authoring.writer()), the operation is:

js
await writer.enableEntityKeys("string");
const first = await writer.createEntity("document-1", {
  text: { kind: "string", value: "first" },
});
await writer.supersedeEntity("document-1", {
  text: { kind: "string", value: "second" },
}, first.revision);

This is an operation fragment, not standalone backend setup. Integer keys use BigInt where needed; bytes and revision tokens use Uint8Array. The full JS setup and pinned-history example and Python public test pin these calling conventions.

Ragged and CSR: components, not a padded dense array

Python 0.0.12:

python
from pathlib import Path
from tempfile import TemporaryDirectory
import numpy as np
import dreamdb as db

schema = db.Schema().add_ragged_array("rows", "f32")
value = {
    "row_offsets": np.array([0, 0, 2], dtype=np.uint64),
    "values": np.array([0.0, 2.5], dtype=np.float32),
}
with TemporaryDirectory(prefix="dreamdb-ragged-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create("ragged", schema, backend=backend)
    ds.append_many([{"_anchor": 7, "rows": value}])
    reopened = db.Dataset.open("ragged", backend=backend)
    got = reopened.get_array_item("rows", 7)
    np.testing.assert_array_equal(got["row_offsets"], value["row_offsets"])
    np.testing.assert_array_equal(got["values"], value["values"])

Here the first row is empty and the second has two values. For a 2 x 4 CSR matrix, declare add_sparse_csr("rows", "f32", (2, 4)) instead, and add column_indices=np.array([1, 3], dtype=np.uint64) to the components. Offsets, column ordering/bounds and declared dtype are validated; do not silently cast a malformed structure into validity. In WASM the corresponding typed component arrays use explicit unsigned-64 indices/offsets, not lossy JavaScript Numbers.

See the public component roundtrip and the array contract. Dense-array examples in released-package documentation do not imply these structured methods exist in those releases.

Geometry: increasing byte budgets return only the new suffix

Python 0.0.12:

python
import struct
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-geometry-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create(
        "scene", db.Schema().add_geometry("points"), backend=backend
    )
    # point record: quantized xyz (3 little-endian u16) + RGBA (4 u8)
    records = b"".join(
        struct.pack("<3H4B", x, 0, 0, 255, 128, 0, 255)
        for x in range(300, -1, -1)
    )
    ds.publish_geometry_cells("points", 42, records)
    ds = db.Dataset.open("scene", backend=backend)
    reader = ds.open_geometry_item("points", 42)
    first = reader.extend_cell_prefix(0, 1000)
    suffix = reader.extend_cell_prefix(0, 2000)
    assert len(first) == len(suffix) == 1000
    accumulated = first + suffix
    assert reader.extend_cell_prefix(0, 500) == b""
    del reader  # release the session/cache when no longer needed

The request budget is a cumulative prefix, not “read another N bytes”. Retain returned suffixes if the application needs the accumulated result. Writer ordering is canonical, not necessarily the input order. Quantization is explicit and lossy; the example uses the default cell frame, not arbitrary world-coordinate floats.

Mesh uses publish_geometry_mesh and read_mesh_lod: each selected LOD is a complete independently decodable unit, not a progressive point prefix. WASM uses openGeometryItem / extendCellPrefix with BigInt anchor/cell/budget arguments and readMeshLod with the checked LOD index. Release wasm-bindgen handles with free() when finished.

v1 has a ≤1 MiB outer Track, ≤1 MiB cells and ≤4 MiB data packs. The producer works in memory; metadata selection may read all hierarchy pages, and session cache grows until released. Mesh error is a producer assertion, not a verified simplification bound. Geometry publication on keyed Datasets and explicit geometry compaction refuse in this version.

See spec/0027, Python example and JS boundary example.

Vector identity, distance and graph queries

For identified embedding fields, carry the producer's identity into every write and query. Main Python append_many accepts the per-field embedding_specs map; append_hot additionally accepts a single spec_id, and iter_vector accepts spec_id. Do not substitute the field's expected identity for the producer's true identity merely to make a check pass. Keep legacy unidentified fields on their explicit legacy contract. Binding implementation.

L2 is declared with algorithm="dreamdb.lsh-l2", preserving input magnitude. A Python 0.0.12 Schema can use add_embedding("v", dim=2, algorithm="dreamdb.lsh-l2"); the public L2 test shows the magnitude-sensitive ordering and reopening. Probe depth still affects ANN candidate recall.

For graphs, build_graph_index establishes the raw graph. compress_graph_index takes compressor bytes and an explicit rerank policy; exact sources must remain available when required. append_graph_nodes and consolidate_graph operate on raw Fresh lineages; they are not a supported incremental compressed-graph path. The public graph examples cover both separately.

Federation router bindings name the exact eligible shard snapshots. Updating a child does not automatically update an old router binding. Do not describe routed top-K as an exhaustive global search or successful partial availability as full coverage. Routing contract.

Maintenance and hybrid-query limits

HotShard append/flush is field-qualified. Scalar/text payloads are typed rather than reinterpreted as vector bytes. flush_hot and consolidate_graph are explicit operations; there is no always-running TTL/consolidation daemon. Compaction, consolidation and GC are separate actions; retained roots keep old data alive.

Main hybrid query planning exposes the chosen plan, executed candidate depths and degradation status. The max_latency_ms budget is interpreted through a cost model, not enforced as a real-time deadline. Preserve predicates and exactness requirements when a budget cannot fund the desired search. Scalar B-tree path-copying reduces update amplification without changing scalar value semantics. See hybrid planning and scalar indexing.

Evidence and release discipline

Examples above are source-matched adaptations of existing public SDK boundary tests at the pinned main commit. They target Python 0.0.12/npm 0.5.4 rather than Python 0.0.10/npm 0.5.2. The corresponding product changes already passed their PR checks. Document/link builds do not re-prove product behavior, WAN performance or cross-implementation agreement.

Release acceptance uses the existing public SDK boundary suites against the installed wheel/tarball. It does not imply every prose example or performance claim was independently tested. Stable version metadata is updated with the SDK page review, not inferred from specification maturity.