# API Reference

> **Version scope:** Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive
> geometry, structured ragged/CSR and additional index capabilities are included
> in these releases. See [feature examples and limits](/main-features.md) for the
> supported boundaries; older packages may not expose these methods.

Selected public signatures are matched to the 0.0.12 release source. This page covers common pipeline operations; the [feature guide](/main-features.md) adds query controls, entity-key, geometry and structured-array examples rather than claiming an exhaustive API inventory.

Looking for SQL? See [SqlSession in the SQL guide](/sql.md#python).
It is an **unreleased source adapter**, not part of the wheel version documented
on this API page; the guide includes the required revision and local build steps.

## Schema

Chainable field declarations. Every method returns the schema.

```python
Schema.add_image(name, mime="jpeg", required=True, chunk_size=None, pack_items=None)
Schema.add_video(name, mime="mp4", required=True, chunk_size=None, pack_items=None)
Schema.add_audio(name, mime="wav", required=True, chunk_size=None, pack_items=None)
Schema.add_embedding(name, dim, algorithm="dreamdb.lsh-cosine", required=True,
                     lsh_bits=None, compressor=None, spatial_index=None, rerank=False,
                     bucket_max_bytes=None, embed_spec=None)
Schema.add_geometry(name, family="point")
Schema.add_ragged_array(name, dtype, inner_shape=(), **kwargs)
Schema.add_sparse_csr(name, dtype, shape, **kwargs)
Schema.add_array(name, dtype, shape, *, endianness="little", layout="c",
                 codec="raw", semantic_type="none", units="none", frame="none",
                 axis_order="none", track_kind="event", object_kind="unbucketed",
                 required=True)
Schema.add_scalar_categorical(name, required=True)
Schema.add_scalar_string(name, required=True)
Schema.add_scalar_int(name, required=True)
Schema.add_scalar_float(name, required=True)
Schema.add_scalar_bool(name, required=True)
Schema.add_scalar_timestamp(name, required=True)
```

`algorithm` defaults to `dreamdb.lsh-cosine`, which requires a published `spatial_index` hash like every other partitioning algorithm — see [Indexes](/python-sdk-indexes.md). `compressor` takes a hash from `publish_rabitq_compressor` or `train_and_publish_pq_compressor`; `rerank=True` keeps full-precision vectors alongside compressed codes for a second scoring pass.

`add_array` declares a fixed-shape numeric value. Python accepts and returns NumPy arrays, validating the persisted dtype, shape, byte order, layout, codec, units, semantic type, frame, and axis order. The supported v1 dtypes are `f32`, `f64`, `i8/i16/i32/i64`, and `u8/u16/u32/u64`.

## Dataset

### Lifecycle

```python
Dataset.create(name: str, schema, backend: str) -> Dataset
Dataset.open(name: str, schema=None, backend: str = "") -> Dataset
Dataset.open_at(version: dict, backend: str) -> Dataset
Dataset.open_ref(ref_name: str, backend: str) -> Dataset
Dataset.open_by_manifest(manifest_hash: str, backend: str) -> Dataset
```

`open` recovers the schema from the manifest; pass one only if you want it checked field-for-field. `open_at` takes the dict `snapshot()` returned.

### Writing

```python
append_many(samples, commit: bool = True, hot: bool = False, embedding_specs=None) -> int
add_array(name, dtype, shape, *, ..., required=False) -> Dataset
add_constant(field_name: str, modality: str, value) -> None
get_constant(field_name: str)
add_array_constant(field_name: str, value, *, dtype, shape, ...) -> None
get_array_constant(field_name: str)
commit() -> None
delete(anchors, reason: str | None = None) -> str
configure_hot_shard(flush_threshold: int = 10_000, ttl_seconds: int = 30) -> None
flush_hot() -> None
ingest_video(field: str, path: str, *, frag_duration: float = 2.0,
             anchor: int = 0, preview: dict | None = None) -> dict
```

`samples` is a list of dicts keyed by field name. `hot=True` routes writes through a hot shard; threshold/age checks run during ingest, not in an independent background timer. Use `flush_hot()` when an explicit flush is needed. `embedding_specs` supplies per-field embedding identities, while vector queries use `spec_id`. `ingest_video` fragments a file with ffmpeg, so ffmpeg has to be on `PATH`.

`Dataset.add_array` adds an optional typed-array field to an existing dataset. `add_constant` supports the SDK's defined text, URI, SPDX, and JSON constant modalities; `add_array_constant` stores one typed NumPy array for the complete timeline.

### Reading and querying

```python
iter_vector(field: str, query, top_k: int, batch_size: int = 64,
            where_eq=None, where=None, as_numpy: bool = False,
            nprobe: int | None = None, fields=None, spec_id=None,
            rerank_pool_size: int | None = None, rerank_mode: str = "inherit")
iter_scalar(where_eq=None, where=None, batch_size: int = 64)
iter_stream(batch_size: int = 256, fields=None, channel_buffer: int = 8,
            start_ns: int = 0, end_ns: int | None = None)
iter_all_batches(...)
iter_arrow_batches(batch_size: int = 256, fields=None,
                   shuffle_seed: int | None = None, window_batches: int | None = 256)
query_scalar(field: str, op: str, value) -> list[int]
query_text(field: str, query: str, top_k: int = 10) -> list[tuple[int, float]]
query_hybrid(...)
query_phrase(...)
get_by_anchor(field: str, root: bytes, anchor: int)
video_item_by_key(field: str, item_key) -> dict | None
video_item_at(field: str, anchor: int) -> dict | None
read_video_item_range(field: str, item_key,
                      relative_start: int, relative_end: int) -> dict | None
distinct_values(field: str) -> list[tuple]
count() -> int
list_anchors() -> list[int]
```

The `iter_*` methods yield dicts of columns plus `_time_anchors`. `query_*` methods return anchors, or `(anchor, score)` tuples for the scored ones. `nprobe` widens a vector search across neighbouring cells.

`rerank_pool_size` is a positive integer at least `top_k`; omitted, it preserves
the five-times-top-k partition pool. `rerank_mode` is `inherit`, `approximate`
or `exact`. Explicit modes override the Schema for this query only; exact mode
refuses missing exact capability rather than falling back. These controls do
not widen probes or guarantee global recall. Graph queries reject explicit pool
size or non-inherit mode. See [query examples and limits](/main-features.md).

The VideoItem metadata methods do not download media. `item_key` is opaque bytes and anchors remain exact Python integers. `read_video_item_range` returns initialization bytes and complete fragments overlapping an item-relative half-open range; it materializes the result in memory and never crosses into the next item.

> **Warning:** **The raw `iter_stream` fast path still requires an embedding field.** A projection without one raises `streaming iter with no embedding fields — use iter() for scalar-only`:
> 
> <!--CODEBLOCK:4-->
> 
> Scalar-only datasets are otherwise usable: `query_scalar`, `iter_scalar`, `distinct_values`, `list_anchors`, `count`, and `history` all work without an embedding.

`count()` now works with every schema and excludes tombstoned anchors. It counts the union of distinct visible anchors across Tracks, so it is `O(U_total)` and materializes that anchor set rather than reading a stored counter.

`iter_arrow_batches` uses the embedding-only stream when exactly one embedding field is requested, and a windowed eager join for general projections. `window_batches=None` restores the fully eager behavior. Typed-array columns use Arrow `FixedShapeTensor`. The method needs `pyarrow` and NumPy, and says so when either is missing:

```
iter_arrow_batches requires `pyarrow`. Install with `pip install pyarrow`
or use `iter_vector` for the plain-dict path.
```

### Versioning

```python
snapshot(label: str) -> dict
branch(new_name: str) -> Dataset
merge(other_ref_name: str, strategy: str = "fast-forward") -> str
merge_many(branches) -> str
history(max_depth: int = 50)
current_manifest()
timeline()
ref_name()
list_refs()
delete_ref(name: str) -> None
tombstone_set() -> list[int]
```

`snapshot` returns `{"label": "<ref>@<label>", "manifest": "<hash>", "timeline": "<hash>"}` — keep it to reopen that exact state. Details in [Versioning](/python-sdk-versioning.md).

### Indexes and layers

```python
add_text_index(field: str, parent_field: str, docs: list[tuple[int, str]],
               k1: float = 1.5, b: float = 0.75, delta: float = 1.0, algorithm: str = ...)
add_embedding_layer(...)
add_image_layer(...)
append_embedding_layer(...)
build_graph_index(...)
build_anchor_index(field: str) -> bytes
add_multi_vector_index(...)
add_splade_index(...)
compact(modality: str | None = None, threshold: int = 1, max_cells: int = 0) -> dict
```

A layer attaches a derived field to an existing one — a second embedding of the same images, a thumbnail, a text index. `compact` returns counters:

```python
ds.compact()
# {'cells_examined': 21, 'cells_compacted': 0, 'fragments_collapsed': 0, …}
```

### Metadata

```python
set_meta(key: str, value: str) -> None
meta() -> dict
set_description(text: str) -> None
description() -> str | None
set_coverage_complete(...)
```

## Module functions

Jobs that happen outside a dataset's lifetime.

```python
train_and_publish_ivf_centroids(backend, sample, dim, k, iterations=10, seed=None) -> bytes
train_and_publish_imi_centroids(backend, sample, dim, k_sub, iterations=10, seed=None) -> bytes
train_and_publish_pq_compressor(backend, sample, dim, m, k, iterations=10, seed=None) -> bytes
publish_rabitq_compressor(backend, dim, seed=None, bits_per_dim=1,
                          with_correction_factors=True) -> bytes
train_and_publish_vamana_graph(backend, sample, dim, modality, r=32, alpha=1.2,
                               l_build=100, ...) -> bytes
gc(backend, keep_manifests=10, keep_since_secs=0, dry_run=True,
   *, roots=None, writers_paused=False) -> None
compare_refs(refs: list[str], fields: list[str], backend: str, batch_size=256)
```

Each `train_and_publish_*` returns a 33-byte content multihash to pass as `spatial_index=` or `compressor=`. They are deterministic: the same sample, parameters, and seed produce the same hash, so two workers publishing independently is a content-addressed no-op rather than a conflict.

`gc` defaults to `dry_run=True`. Destructive GC additionally requires
`writers_paused=True` after externally stopping and draining writers, uploads,
staged transactions and root changes, with a settled Ref listing. Keep that
fence for the whole run; the flag acquires no lock. Use `roots` (base32 Manifest
hashes) to retain otherwise unrooted reader snapshots. An empty root set refuses.
A concurrent dry-run is advisory, not a reusable deletion plan; age retention
cannot protect a concurrent writer reusing old Objects. Version 0.0.12 fixes
reference-mode vector-pool retention, but does not recover already-deleted data.

## PyTorch

```python
from dreamdb.torch import DreamDBIterableDataset

DreamDBIterableDataset(dataset, *, field, query, top_k, batch_size=64,
                       where_eq=None, as_numpy=True, transform=None)
```

See [Training pipelines](/python-sdk-training.md).
