# Feature guide — Python 0.0.12 / npm 0.5.4

This guide targets Python 0.0.12 and npm 0.5.4. The core baseline is
[a119479](https://github.com/dreamlake-ai/dreamdb-core/commit/a119479d6a3a2d347081e15cf1e573f26a48c011).
The entity, geometry and array examples originated in the preceding release;
the new query controls and GC repair are described below. Older packages do not
necessarily expose these options. Rust-only operator capabilities are called out
separately, not presented as Python or JavaScript methods.

For **read-only SQL v1**, see the [SQL guide](/sql.md): merged Rust API plus
unreleased Python, npm and CLI source adapters. These are separate from the
SDK releases described here; a source implementation is not a package release.

## Upgrade collectors before reference-mode GC

Version 0.0.12 includes the fix for GC deleting VectorStorage referenced only
through reference-mode Bucket records (#365). Both inline and paged indexes now
use authenticated dependency traversal; malformed metadata refuses before sweep.
Upgrade or fence **every collector** processing these roots. Updating only a
reader does not make an old collector safe. The repair cannot recover previously
deleted Objects or repair #89's historical Manifest metadata damage.

GC still requires an externally enforced maintenance window and retained reader
roots. An age threshold and a prior dry-run do not substitute for that fence.
See [GC requirements](/spec-protocol-operations.md#732-retention-threshold-and-maintenance-precondition).

## Query-local exactness and pool size

For an existing partition-vector field and a query of the declared dimension:

```python
batches = ds.iter_vector(
    "embedding", query, top_k=10,
    rerank_pool_size=50, rerank_mode="exact",
)
```

```js
const hits = await space.queryVector("embedding", query, {
  topK: 10, rerankPoolSize: 50, rerankMode: "exact",
});
```

These are call fragments; use the [Python](/python-sdk-quickstart.md) or
[JavaScript](/typescript-sdk-quickstart.md) quickstart to create/open the dataset.
`inherit` keeps the Schema policy; `approximate` skips the exact rerank pass;
`exact` requires an exact source for every selected candidate, with no silent
fallback. Pool size is a positive integer at least top-k (default five times
top-k), selected after eligibility filtering. It does not widen probe coverage
or guarantee global exact recall. Graph queries reject an explicit pool or
non-inherit mode. Raw L2 vectors retain their exact metric scores in all modes.

## Rust operator APIs: not new SDK wrappers

The source release also includes fragmentation observations, explicit maintenance
plans, bounded Fragment repacking, resumable Reencode admission, pull-based Ref
Watch, create-only snapshot tags, media Stream selection and verified closure-copy
plans. These are Rust Dataset/operator APIs; no corresponding new Python or
JavaScript wrappers are implied. See [release notes](/release-notes.md) and the
owning [operations](/spec-protocol-operations.md), [evolution](/spec-schema-evolution.md)
and [copy-plan](/spec-federation.md) contracts. None adds a resident maintenance
scheduler or an automatic business rollout.

## What changed and what did not

| Capability | Main behavior | Important boundary |
| --- | --- | --- |
| Entity keys (#322) | Optional string/bytes/int keys, stable entity ID, revision-checked atomic create/update/delete | Timeline-scoped; v1 key index ≤16 MiB; divergent key-index merges refuse even for disjoint edits |
| Geometry (#316) | Point/splat cells, packed data, paged hierarchy, independent mesh LODs, authenticated progressive reads | One outer Item, not one Sample per point; metadata/cache and producer memory are not universally bounded |
| Ragged and CSR (#323) | Dataset/Python/WASM writers and typed reads, not just codec declarations | Structured Arrow projection unsupported; dense/ragged/sparse remain different declared families |
| L2 (#321) | Magnitude-preserving Euclidean partition/search path | Do not normalize inputs as cosine vectors or interpret its scores as cosine similarity |
| Compressed graphs (#320) | ADC traversal with exact-source pool reranking | Exact reranking cannot recover nodes excluded from candidate selection |
| Federation routing (#319) | Exact-bound shard router over explicitly eligible children | Does not discover backends or guarantee global exact recall |
| Fresh graphs (#318) | Append plus id-preserving consolidation; changed pages published, others reused | Raw graph lineage only; explicit calls, no resident scheduler or physical erasure |
| HotShards (#324/#328) | Field-qualified hot embedding, scalar and text data | Flush age checked by ingest is not an independent timer; configure before depending on hot behavior |
| Embedding identity (#317, PR #327) | spec_id checked at append, query and index-build boundaries | Approximate compatibility never authorizes reusing a vector from another identity |
| Hybrid planning (#326) | Executed plan, configurable latency estimate/budget, observable trace | Cost estimates are not a hard wall-clock SLA; no silent unfiltered fallback |
| Scalar B-trees (#325) | Persistent path-copy updates | Does not mean all indexes or all maintenance operations are constant-memory |

The same Rust implementation underlies the SDKs. Shared-code tests are not
independent implementations agreeing on a protocol.

## Entity keys: keep the revision token

Python 0.0.12:

```python
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-keys-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create(
        "keys", db.Schema().add_scalar_string("text"), backend=backend
    )
    ds.enable_entity_keys("string")
    first = ds.create_entity("document-1", {"text": "first"})
    second = ds.supersede_entity(
        "document-1", {"text": "second"}, expected=first["revision"]
    )
    assert second["entity_id"] == first["entity_id"]
    reopened = db.Dataset.open("keys", backend=backend)
    assert reopened.get_entity("document-1")["sample"] == {"text": "second"}
```

A stale expected revision must refuse, not overwrite the newer value. Keep the
returned token as bytes; do not replace it with a timestamp or decoded text.
Logical deletion leaves historical snapshots readable. Deleted entries can be
explicitly restored using their latest revision; an ordinary scalar filter does
not provide any of these transaction guarantees. Read
[spec/0026](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0026-entity-keys.md) before enabling keyed writes.

For a WASM 0.5.4 writer, the corresponding calls are
enableEntityKeys, createEntity, supersedeEntity, deleteEntity and upsertEntity.
An existing Space is pinned: reopen to read a later publication. Given an
initialized Writer named writer (from Authoring.writer()), the operation is:

```js
await writer.enableEntityKeys("string");
const first = await writer.createEntity("document-1", {
  text: { kind: "string", value: "first" },
});
await writer.supersedeEntity("document-1", {
  text: { kind: "string", value: "second" },
}, first.revision);
```

This is an operation fragment, not standalone backend setup. Integer keys use
BigInt where needed; bytes and revision tokens use Uint8Array. The full
[JS setup and pinned-history example](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-wasm/test/js_boundary.mjs#L392)
and [Python public test](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/tests/test_entity_keys.py)
pin these calling conventions.

## Ragged and CSR: components, not a padded dense array

Python 0.0.12:

```python
from pathlib import Path
from tempfile import TemporaryDirectory
import numpy as np
import dreamdb as db

schema = db.Schema().add_ragged_array("rows", "f32")
value = {
    "row_offsets": np.array([0, 0, 2], dtype=np.uint64),
    "values": np.array([0.0, 2.5], dtype=np.float32),
}
with TemporaryDirectory(prefix="dreamdb-ragged-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create("ragged", schema, backend=backend)
    ds.append_many([{"_anchor": 7, "rows": value}])
    reopened = db.Dataset.open("ragged", backend=backend)
    got = reopened.get_array_item("rows", 7)
    np.testing.assert_array_equal(got["row_offsets"], value["row_offsets"])
    np.testing.assert_array_equal(got["values"], value["values"])
```

Here the first row is empty and the second has two values. For a 2 x 4 CSR
matrix, declare add_sparse_csr("rows", "f32", (2, 4)) instead, and add
column_indices=np.array([1, 3], dtype=np.uint64) to the components. Offsets,
column ordering/bounds and declared dtype are validated; do not silently cast
a malformed structure into validity. In WASM the corresponding typed component
arrays use explicit unsigned-64 indices/offsets, not lossy JavaScript Numbers.

See the [public component roundtrip](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/tests/test_typed_arrays.py#L15)
and [the array contract](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0025-typed-array-items.md).
Dense-array examples in released-package documentation do not imply these
structured methods exist in those releases.

## Geometry: increasing byte budgets return only the new suffix

Python 0.0.12:

```python
import struct
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-geometry-") as directory:
    backend = Path(directory).as_uri()
    ds = db.Dataset.create(
        "scene", db.Schema().add_geometry("points"), backend=backend
    )
    # point record: quantized xyz (3 little-endian u16) + RGBA (4 u8)
    records = b"".join(
        struct.pack("<3H4B", x, 0, 0, 255, 128, 0, 255)
        for x in range(300, -1, -1)
    )
    ds.publish_geometry_cells("points", 42, records)
    ds = db.Dataset.open("scene", backend=backend)
    reader = ds.open_geometry_item("points", 42)
    first = reader.extend_cell_prefix(0, 1000)
    suffix = reader.extend_cell_prefix(0, 2000)
    assert len(first) == len(suffix) == 1000
    accumulated = first + suffix
    assert reader.extend_cell_prefix(0, 500) == b""
    del reader  # release the session/cache when no longer needed
```

The request budget is a cumulative prefix, not “read another N bytes”.
Retain returned suffixes if the application needs the accumulated result.
Writer ordering is canonical, not necessarily the input order. Quantization is
explicit and lossy; the example uses the default cell frame, not arbitrary
world-coordinate floats.

Mesh uses publish_geometry_mesh and read_mesh_lod: each selected LOD is a
complete independently decodable unit, not a progressive point prefix.
WASM uses openGeometryItem / extendCellPrefix with BigInt anchor/cell/budget
arguments and readMeshLod with the checked LOD index. Release wasm-bindgen
handles with free() when finished.

v1 has a ≤1 MiB outer Track, ≤1 MiB cells and ≤4 MiB data packs. The producer
works in memory; metadata selection may read all hierarchy pages, and session
cache grows until released. Mesh error is a producer assertion, not a verified
simplification bound. Geometry publication on keyed Datasets and explicit
geometry compaction refuse in this version.

See [spec/0027](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0027-progressive-geometry.md),
[Python example](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/tests/test_geometry.py) and
[JS boundary example](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-wasm/test/js_boundary.mjs#L420).

## Vector identity, distance and graph queries

For identified embedding fields, carry the producer's identity into every write
and query. Main Python append_many accepts the per-field embedding_specs map;
append_hot additionally accepts a single spec_id, and iter_vector accepts spec_id.
Do not substitute the field's
expected identity for the producer's true identity merely to make a check pass.
Keep legacy unidentified fields on their explicit legacy contract.
[Binding implementation](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/src/lib.rs#L924).

L2 is declared with algorithm="dreamdb.lsh-l2", preserving input magnitude.
A Python 0.0.12 Schema can use
add_embedding("v", dim=2, algorithm="dreamdb.lsh-l2"); the
[public L2 test](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/tests/test_l2_index.py) shows the
magnitude-sensitive ordering and reopening. Probe depth still affects ANN
candidate recall.

For graphs, build_graph_index establishes the raw graph. compress_graph_index
takes compressor bytes and an explicit rerank policy; exact sources must remain
available when required. append_graph_nodes and consolidate_graph operate on
raw Fresh lineages; they are not a supported incremental compressed-graph path.
The [public graph examples](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/dreamdb-dataset-python/tests/test_compressed_graph.py)
cover both separately.

Federation router bindings name the exact eligible shard snapshots. Updating a
child does not automatically update an old router binding. Do not describe
routed top-K as an exhaustive global search or successful partial availability
as full coverage. [Routing contract](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0013-graph-indexing.md#6-routing-graphindex--cross-shard-knn-resolves-oq-51).

## Maintenance and hybrid-query limits

HotShard append/flush is field-qualified. Scalar/text payloads are typed rather
than reinterpreted as vector bytes. flush_hot and consolidate_graph are explicit
operations; there is no always-running TTL/consolidation daemon. Compaction,
consolidation and GC are separate actions; retained roots keep old data alive.

Main hybrid query planning exposes the chosen plan, executed candidate depths
and degradation status. The max_latency_ms budget is interpreted through a cost
model, not enforced as a real-time deadline. Preserve predicates and exactness
requirements when a budget cannot fund the desired search. Scalar B-tree
path-copying reduces update amplification without changing scalar value
semantics. See [hybrid planning](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0015-hybrid-retrieval.md) and
[scalar indexing](https://github.com/dreamlake-ai/dreamdb-core/blob/65c8e05d32a8774124a4a3c6f97ba57c47bff4a8/spec/0011-scalar-indexing.md).

## Evidence and release discipline

Examples above are source-matched adaptations of existing public SDK boundary
tests at the pinned main commit. They target Python 0.0.12/npm 0.5.4 rather
than Python 0.0.10/npm 0.5.2. The corresponding product changes
already passed their PR checks. Document/link builds do not re-prove product
behavior, WAN performance or cross-implementation agreement.

Release acceptance uses the existing public SDK boundary suites against the
installed wheel/tarball. It does not imply every prose example or performance
claim was independently tested. Stable version metadata is updated with the
SDK page review, not inferred from specification maturity.
