# Quickstart

> **Version scope:** Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive
> geometry, structured ragged/CSR and additional index capabilities are included
> in these releases. See [feature examples and limits](/main-features.md) for the
> supported boundaries; older packages may not expose these methods.

Forty records with embeddings, labels, and captions, written to a directory and queried back. No services, no cloud account.

```bash
pip install dreamdb numpy
mkdir -p /tmp/notes
```

## Train, declare, create

A partitioning index decides where each vector is stored, so it has to exist before the first record does. `dreamdb.ivf-cosine` needs a sample of representative vectors to train on — a few hundred is enough to start, and `k ≈ √N` is the rule of thumb for the production corpus size:

```python
import numpy as np
import dreamdb as db

BACKEND = "file:///tmp/notes"
rng = np.random.default_rng(5)

index = db.train_and_publish_ivf_centroids(
    BACKEND, rng.standard_normal((512, 32)).astype(np.float32), dim=32, k=16,
)

schema = (db.Schema()
          .add_embedding("embedding", dim=32,
                         algorithm="dreamdb.ivf-cosine", spatial_index=index)
          .add_scalar_categorical("label")
          .add_scalar_string("caption"))

ds = db.Dataset.create("notes", schema, backend=BACKEND)
```

## Append

One dict per record, keyed by field name. `numpy` arrays are accepted directly:

```python
captions = ["a brown dog on grass", "a tabby cat asleep",
            "a red car at night", "two dogs running"]

ds.append_many([
    {"embedding": rng.standard_normal(32).astype(np.float32),
     "label": ["cat", "dog"][i % 2],
     "caption": captions[i % 4]}
    for i in range(40)
])

print(ds.count())                      # 40
print(ds.distinct_values("label"))     # [('cat', 20), ('dog', 20)]
```

`append_many` commits by default. Pass `commit=False` to batch several appends into one manifest and call `commit()` when you are done.

## Query

Vector search returns batches, not rows — the reader streams whole buckets:

```python
query = rng.standard_normal(32).astype(np.float32)

for batch in ds.iter_vector(field="embedding", query=query, top_k=5, batch_size=5):
    print(batch["label"], batch["_time_anchors"][:1])
```

Every batch is a dict of columns plus `_time_anchors`. Add `as_numpy=True` to get the embedding column as an `ndarray` instead of a list.

Filters compose with the vector search, so the scan never widens:

```python
hits = ds.iter_vector(field="embedding", query=query, top_k=5,
                      where_eq={"label": "dog"}, as_numpy=True)
```

Scalar-only iteration skips vectors entirely:

```python
rows = sum(len(b["_time_anchors"]) for b in ds.iter_scalar(where_eq={"label": "cat"}))
print(rows)     # 20
```

## Reopen it later

The schema is recovered from the manifest, so it does not need re-declaring:

```python
ds = db.Dataset.open("notes", backend=BACKEND)
```

## Full-text search

Captions are searchable once a BM25 index exists over them. Build it from `(anchor, text)` pairs:

```python
anchors = ds.list_anchors()
docs = [(a, captions[i % 4]) for i, a in enumerate(anchors)]

ds.add_text_index("caption_idx", "caption", docs)

print(ds.query_text("caption_idx", "dog", 3))
# [(1785381710580097000, 1.2619806528091431),
#  (1785381710580101000, 1.2619806528091431),
#  (1785381710580105000, 1.2619806528091431)]
```

`query_text` returns `(anchor, score)` tuples. The tokenizer is built in, so no model is involved — [more on indexes](/python-sdk-indexes.md#text-indexes).

## Next

- **[Versioning](/python-sdk-versioning.md)** — Snapshots, branches, sharded ingest, and deletion.

- **[Indexes](/python-sdk-indexes.md)** — Which index to train, and which compressor to pair with it.

- **[Read it from an app](/typescript-sdk-quickstart.md)** — Serve the directory and query it from JavaScript.
