Modeling your data
Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.
Most of the work of putting a dataset into DreamDB is deciding which field type each artifact gets. The important distinction is between data you search by similarity and numeric arrays you store with an exact declared shape.
What has a field type
| Kind | Declare with | Stores |
|---|---|---|
| Media | add_image, add_video, add_audio | Opaque bytes with a MIME type |
| Vectors | add_embedding | A fixed-dimension float32 vector, indexed for search |
| Typed arrays | add_array | A fixed-shape numeric value with declared dtype, layout, units, and frame |
| Scalars | add_scalar_categorical, _string, _int, _float, _bool, _timestamp | One value per record, indexable and filterable |
| Constants | add_constant, add_array_constant | One value for the complete timeline |
Python 0.0.12 supports dense arrays and separate ragged-array/sparse-CSR families; see structured-array examples. Structured ragged/CSR Arrow projection is unsupported. Strided and device-native forms remain deferred. Structs are still modeled as several named fields.
When your data is an array
Per-frame poses, mesh vertices, point clouds, joint angles, and calibration matrices use a typed-array field:
DreamDB persists and validates the dtype, shape, endianness, C/Fortran layout, codec, semantic type, units, frame, and axis order. Python writes and reads NumPy arrays. Raw and NPY codecs are supported for f32, f64, and signed or unsigned 8/16/32/64-bit integers.
Not an embedding. add_embedding looks array-shaped, but it is an index type, not a storage type. The cosine algorithms need a direction, so an all-zero vector is refused outright:
That is correct behaviour for a cosine index — a zero vector has no direction to bucket by. It just means embedding fields are for things you intend to search by similarity, not for arbitrary numeric arrays.
Do not use add_embedding as generic numeric storage. Cosine embeddings are normalized for similarity search; the L2 path preserves magnitude. Typed arrays preserve the declared numeric value without imposing either search metric. Neither is a substitute for the other.
Constants that apply to a whole dataset
Use a Constant when one value applies to the whole timeline:
Generic Python constants have defined codecs for title, author, license, description, source URI, and author JSON. Typed-array constants return NumPy arrays and need no invented anchor.
Anchors are the primary record key
For ordinary Track records, the time anchor is the primary lookup and join key. Two records written at the same anchor are, as far as anything anchor-keyed is concerned, one:
Both rows can be stored and both scalar values can survive. What collapses is addressability: one anchor now denotes two records, so count(), joins, and point reads treat them as one logical position.
The fix is to choose anchors deliberately rather than let two logical rows land on the same instant. Set them explicitly with the reserved _anchor key (nanoseconds since the epoch), and offset dataset-level rows away from frame timestamps.
This is why spec 0000 now describes the timeline as the primary index axis rather than the only key, and notes that entities whose identity is independent of time carry an explicit id field. If your data has rows that are not naturally time-anchored, model that id yourself — as a scalar field — rather than relying on the anchor to distinguish them. VideoItem is the built-in exception: a logical video item has an opaque stable byte key in addition to its timeline extent. Use video_item_by_key for identity lookup and video_item_at for timeline lookup.
A worked shape
A per-clip capture with per-frame data ends up looking like:
Search by the embedding, filter by the scalars, and receive pose as a validated NumPy array.