# Modeling your data

> **Version scope:** Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive
> geometry, structured ragged/CSR and additional index capabilities are included
> in these releases. See [feature examples and limits](/main-features.md) for the
> supported boundaries; older packages may not expose these methods.

Most of the work of putting a dataset into DreamDB is deciding which field type each artifact gets. The important distinction is between data you search by similarity and numeric arrays you store with an exact declared shape.

## What has a field type

| Kind | Declare with | Stores |
|---|---|---|
| Media | `add_image`, `add_video`, `add_audio` | Opaque bytes with a MIME type |
| Vectors | `add_embedding` | A fixed-dimension float32 vector, indexed for search |
| Typed arrays | `add_array` | A fixed-shape numeric value with declared dtype, layout, units, and frame |
| Scalars | `add_scalar_categorical`, `_string`, `_int`, `_float`, `_bool`, `_timestamp` | One value per record, indexable and filterable |
| Constants | `add_constant`, `add_array_constant` | One value for the complete timeline |

Python 0.0.12 supports dense arrays and separate ragged-array/sparse-CSR families; see [structured-array examples](/main-features.md). Structured ragged/CSR Arrow projection is unsupported. Strided and device-native forms remain deferred. Structs are still modeled as several named fields.

## When your data is an array

Per-frame poses, mesh vertices, point clouds, joint angles, and calibration matrices use a typed-array field:

```python
import numpy as np

schema = db.Schema().add_array(
    "pose",
    dtype="f32",
    shape=[17, 3],
    semantic_type="pose",
    units="metres",
    frame="opencv_camera",
    axis_order="xyz",
)

ds.append_many([{"pose": pose.astype(np.float32), ...}])
```

DreamDB persists and validates the dtype, shape, endianness, C/Fortran layout, codec, semantic type, units, frame, and axis order. Python writes and reads NumPy arrays. Raw and NPY codecs are supported for `f32`, `f64`, and signed or unsigned 8/16/32/64-bit integers.

**Not an embedding.** `add_embedding` looks array-shaped, but it is an *index* type, not a storage type. The cosine algorithms need a direction, so an all-zero vector is refused outright:

```
protocol: hash_vector: ivf: vector norm is zero (cannot normalize)
```

That is correct behaviour for a cosine index — a zero vector has no direction to bucket by. It just means embedding fields are for things you intend to search by similarity, not for arbitrary numeric arrays.

> **Warning:** Do not use `add_embedding` as generic numeric storage. Cosine embeddings are normalized for similarity search; the L2 path preserves magnitude. Typed arrays preserve the declared numeric value without imposing either search metric. Neither is a substitute for the other.

## Constants that apply to a whole dataset

Use a Constant when one value applies to the whole timeline:

```python
ds.add_constant("title", "title.text", "Factory capture 17")
ds.add_constant("license", "license.spdx", "CC-BY-4.0")

ds.add_array_constant(
    "camera_calibration",
    calibration,
    dtype="f64",
    shape=[4, 4],
    semantic_type="calibration",
    frame="opencv_camera",
)
calibration = ds.get_array_constant("camera_calibration")
```

Generic Python constants have defined codecs for title, author, license, description, source URI, and author JSON. Typed-array constants return NumPy arrays and need no invented anchor.

## Anchors are the primary record key

For ordinary Track records, the time anchor is the primary lookup and join key. Two records written at the same anchor are, as far as anything anchor-keyed is concerned, one:

```python
A = 1_700_000_000_000_000_000
ds.append_many([
    {"_anchor": A, "clip_id": "clip-1", "frame": -1, ...},   # the clip-level row
    {"_anchor": A, "clip_id": "clip-1", "frame": 0,  ...},   # the first frame
])

ds.count()                    # 1 — count() counts distinct visible anchors
len(set(ds.list_anchors()))   # 1
ds.distinct_values("frame")   # [(0, 1), (-1, 1)]  — both values are there
```

Both rows can be stored and both scalar values can survive. What collapses is addressability: one anchor now denotes two records, so `count()`, joins, and point reads treat them as one logical position.

The fix is to choose anchors deliberately rather than let two logical rows land on the same instant. Set them explicitly with the reserved `_anchor` key (nanoseconds since the epoch), and offset dataset-level rows away from frame timestamps.

> **Note:** This is why spec `0000` now describes the timeline as the *primary index axis* rather than the only key, and notes that entities whose identity is independent of time carry an explicit id field. If your data has rows that are not naturally time-anchored, model that id yourself — as a scalar field — rather than relying on the anchor to distinguish them. VideoItem is the built-in exception: a logical video item has an opaque stable byte key in addition to its timeline extent. Use `video_item_by_key` for identity lookup and `video_item_at` for timeline lookup.

## A worked shape

A per-clip capture with per-frame data ends up looking like:

```python
schema = (db.Schema()
          .add_video("source", mime="mp4")            # the clip
          .add_array("pose", dtype="f32", shape=[17, 3],
                     semantic_type="pose", units="metres",
                     frame="opencv_camera", axis_order="xyz")
          .add_embedding("clip_emb", dim=512, algorithm="dreamdb.ivf-cosine",
                         spatial_index=index)          # what you search by
          .add_scalar_categorical("clip_id")           # the identity anchors don't give you
          .add_scalar_int("frame_idx")
          .add_scalar_float("scale"))
```

Search by the embedding, filter by the scalars, and receive `pose` as a validated NumPy array.
