# Python SDK

> **Version scope:** Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive
> geometry, structured ragged/CSR and additional index capabilities are included
> in these releases. See [feature examples and limits](/main-features.md) for the
> supported boundaries; older packages may not expose these methods.

`dreamdb` is where data gets in. It creates datasets, trains spatial indexes, ingests video, and feeds PyTorch — the parts of DreamDB the browser SDK deliberately doesn't carry.

```bash
pip install dreamdb numpy
```

It is a compiled extension around the same Rust core as the [JavaScript SDK](/typescript-sdk.md), so the two write identical bytes. `numpy` is optional but every example uses it.

## Two classes

- **[Schema](/python-sdk-api.md#schema)** — Declare media, embeddings, fixed-shape numeric arrays, and six scalar types. Chainable.

- **[Dataset](/python-sdk-api.md#dataset)** — Everything else — create, append, query, snapshot, branch, merge, compact, delete.

Module-level functions cover the jobs that happen outside a dataset's lifetime: [training indexes and compressors](/python-sdk-indexes.md), garbage collection, and comparing refs.

## What only Python can do

| | Python | JavaScript |
|---|---|---|
| Create a dataset | Anywhere | Node only |
| Train a spatial index | Yes | No |
| `dreamdb.ivf-cosine` / `imi-cosine` fields | Yes | Cannot be created |
| Build a text (BM25) index | Yes | No |
| Video ingest with fragmenting | Yes, via ffmpeg | No |
| Typed numeric array fields and constants | Yes | Yes, from 0.5.2 |
| Read VideoItems by stable key or time | Yes | Yes, from 0.5.2 |
| PyTorch / Arrow batches | Yes | No |

If a pipeline writes data or builds an index, it wants this SDK. If an application reads and searches, either will do.

## A local directory is a real backend

Nothing needs to be running to use DreamDB from Python:

```python
ds = db.Dataset.create("photos", schema, backend="file:///tmp/photos")
```

`memory://` exists for tests, and `http(s)://` for any S3-compatible endpoint. See [Storage backends](/storage.md).

## Where to go next

- [Quickstart](/python-sdk-quickstart.md) — a dataset, 40 records, and a query
- [Modeling your data](/python-sdk-modeling.md) — which field type each artifact gets, and what to do when none fits
- [API reference](/python-sdk-api.md) — `Schema`, `Dataset`, and the module functions
- [Indexes and compressors](/python-sdk-indexes.md) — what to train, and when
- [Versioning](/python-sdk-versioning.md) — snapshots, branches, sharded ingest, deletion
- [Training pipelines](/python-sdk-training.md) — PyTorch and Arrow
- [Python SDK changelog](/python-sdk-changelog.md) — user-facing changes in each PyPI release
