DreamDB

Spec 0025 — Typed Array Items

Status: Accepted. The reference protocol, Dataset, Python SDK and WASM/TS SDK implement the CLOSED dense-v1 type document, complete identity, registry binding, all three legal Track/storage pairs, raw/npy validation, per-Item and Constant storage, and typed decoding. §11 additionally defines separately versioned ragged-array and sparse-CSR formats with portable conformance vectors. The reference Dataset and Python/WASM SDKs now read and write both families. Depends on: spec/0001, spec/0002, spec/0007, spec/0009, spec/0017, spec/0024. Motivation: A Track can say where its Items are and, after 0002 §5.4, what kind of index addresses them. It cannot say what an Item is. A per-frame 6-DoF pose, a hand's vertex array, a gravity vector and a topology table are dense numeric arrays, and the protocol has no way to declare their dtype, shape, byte order or layout — so they travel as image Items with an invented MIME type and an application-level convention inside the payload. DreamDB can find those bytes and cannot verify or decode them; two conforming readers can disagree about the same Item and neither is wrong. This document defines a declared encoding for dense numeric Items so that any conforming reader decodes them identically, without MIME inference, magic-byte sniffing or an out-of-band schema.


1. Purpose

After 0002 §5.4 and the contextual-decode work, a reader resolving a namespaced modality obtains a TrackKind and an ObjectKind. That is enough to walk the Track's object_index and fetch an Item's bytes. It is not enough to interpret them.

This document adds one thing: a canonical item-type declaration for dense numeric arrays, bound into the modality string so that two arrays of different type are different modalities, therefore different Tracks.

Non-goals for dense v1: ragged and variable-shape arrays, sparse tensors, compressed or strided layouts, and device-native memory formats. §11 defines the first two as new families without changing dense-v1 bytes or meaning.

2. Why the type cannot live only in the registry

The Manifest's registry is keyed by the modality tag string. A writer emits exactly one type entry for one array modality. A reader follows 0002 §5.4's existing rule: byte-identical duplicate entries collapse to one answer, while differing entries under one modality are ambiguous and MUST be rejected.

A single dataset routinely holds several dense arrays of different shape — a [4, 4] pose, a [778, 3] vertex array, a [21, 3] joint array, a [3] gravity vector. If those Tracks share a modality string, the registry can describe at most one of them, and a merge of two Manifests that describe it differently is a registration conflict (0002 §5.4).

Note the converse, settled by issue #41: several fields MAY share one modality when that modality carries only a type declaration and binds no index. Sharing is refused exactly when the shared modality binds an index (spatial_index, scalar_index_hash, text_index, graph_index), because the registry is keyed by modality and the fields would then resolve to the same index. So this document's requirement is precise: distinct item types MUST be distinct modalities; identical item types MAY be shared by any number of fields. That is what makes typed arrays usable — a corpus with fifty [4, 4] f32 pose fields needs one type, not fifty.

This is the same error as binding an index's identity to a modality: a type name is being used as an instance key. The fix is the one 0024 applies to embeddings — the identity-bearing configuration belongs in the modality, so distinct types are distinct modalities and the registry key separates them for free.

3. The item-type document

A canonical CBOR map. Every field is hard: it participates in the digest, and changing any of them produces a different type.

FieldTypeNotes
familytext"dense-array" in v1. Reserved for future families
versionuint1
dtypetextf32, f64, i8, i16, i32, i64, u8, u16, u32, u64
shapearray of uintFixed extents, outermost first. [] means a scalar
endiannesstextlittle or big
layouttextc (row-major) or f (column-major)
codectextraw or npy (§7)
codec_versionuint1
semantic_typetextWhat the numbers mean, e.g. pose, vertices, acceleration
unitstexte.g. metres, radians, none
frametextCoordinate frame, e.g. opencv_camera, world. none if not spatial
axis_ordertexte.g. xyz, rowmajor_rc. none if not applicable

There are no soft fields in v1. Anything that does not affect decoding or meaning — a description, a producer name, a human label — belongs in Manifest metadata, and neither writer nor reader may change behaviour because of it.

The v1 document is CLOSED under 0002 §3.1.4. It contains exactly the twelve keys above, each exactly once. An unknown, missing, repeated or wrongly typed key is a protocol error. Extension uses a new family value governed by a later specification; an implementation that knows only dense-array rejects any other family rather than dropping fields it does not understand.

semantic_type, units, frame and axis_order are hard for the reason 0024 gives for embeddings: a [4, 4] f32 in camera coordinates and a [4, 4] f32 in world coordinates decode identically, validate identically, compose silently, and mean different things. Dimension is not identity, and neither is shape.

3.1 The digest

Encode the document as deterministic CBOR (0002 §3), hash with BLAKE3, and wrap as a 33-byte multihash (0009 §3.1). The base32 form of that multihash is the item type's identity.

The digest is not truncated. Like the corrected 0024 §5 identity, it is the complete 33-byte BLAKE3-256 multihash in canonical 53-character lowercase, unpadded base32 form. A shortened or naked digest cannot identify the declaration and MUST be rejected.

3.2 Manifest registration

The declaration is carried inline in the modality's registry entry under the string key item_type:

"array.f32.kind=event.storage=unbucketed.shape=4x4.item=<multihash>": {
  "item_type": { ... the CLOSED §3 map ... }
}

item_type is the map itself, not encoded CBOR bytes and not an Object address. Its deterministic CBOR encoding is what the modality's item= multihash names. The enclosing registry-entry map remains open under 0002 §3.1.3 so later, independently specified bindings can coexist; item_type itself MUST occur exactly once. Because the declaration is inline and the tag carries its identity, this addition creates no new GC reference or storage lookup.

The tag projection, full identity and registry declaration form one check. A reader MUST NOT accept an array Track from the tag alone, and MUST NOT treat a registry entry without item_type as an untyped array.

4. Modality binding

array.<dtype>.kind=<track-kind>.storage=<object-kind>.shape=<extents>.item=<multihash>

Canonical parameter order is kind, storage, shape, item. Writers MUST emit exactly this order; readers MUST NOT reorder before comparing.

array.f32.kind=event.storage=unbucketed.shape=778x3.item=dz7o3wkkj4k4vmqx…
array.f32.kind=constant.storage=constant.shape=3.item=d2wn2nmxvqlmt7j6…

array is added to the built-in class table of 0002 §5.2. Without it the parser rejects array.* — with AmbiguousNamespace ("cannot locate class segment"), not, as an earlier draft of this section stated, UserDefinedRequiresReverseDns. The class scan finds no known class and no ≥3-segment reverse-DNS prefix, so it cannot decide which segment is the class at all. Verified against ModalityTag::parse (see §14).

Adding the class name is necessary and not sufficient. Every built-in class today derives its TrackKind from the class name alone, and its ObjectKind from the class plus a class-specific flag — embedding reads the bucketed flag, transcript reads bucket. No class derives either kind from a declared parameter. array would be the first, because §4.1 makes all three kind/storage pairings legal for one class.

A conforming implementation MUST therefore extend track_kind() and object_kind() to read the kind and storage parameters for array, and its typed-array resolver MUST reject a tag that omits either — a default would silently pick a Track shape the writer did not declare. The generic modality parser establishes grammar; semantic completeness belongs to typed-array resolution. This is a resolver change, not only a class-table entry.

  • kind = continuous | event | constant determines the TrackKind.
  • storage = unbucketed | time_batch | constant determines the ObjectKind.
  • shape is the extents joined by x, e.g. 778x3. A scalar is shape=0d.
  • item is the full base32 multihash from §3.1.

The modality alphabet is [A-Za-z0-9_] per segment (0002 §5.1), which admits 778x3 and base32 but not hyphens or slashes — which is why frame and units are carried in the hashed document rather than spelled in the tag. A full tag runs about 130 bytes, well inside the 256-byte limit.

4.0 storage values are not ObjectKind wire names

The alphabet has a consequence this document must state rather than leave to be discovered. ObjectKind's wire name for a time-bucketed batch is time-batch, with a hyphen — and a hyphen is not a legal parameter value, so storage=time-batch does not parse. The legal tag spelling is time_batch, which is not an ObjectKind wire name: ObjectKind::from_wire("time_batch") returns None. Both halves verified (§14).

So the mapping is explicit and one-way:

storage (tag)ObjectKind (wire)
unbucketedunbucketed
time_batchtime-batch
constantconstant

A reader MUST translate through this table and MUST NOT pass a storage value to ObjectKind::from_wire directly. A writer MUST NOT emit a storage value outside the left column. Without this, an implementer either emits a tag that does not parse or resolves no ObjectKind at all — and the failure is at the boundary between two specs, where neither side looks wrong on its own.

kindstorageUse
eventunbucketedPer-record arrays at irregular anchors
continuoustime_batchPer-frame arrays batched by time window
constantconstantOne value for the whole Track — calibration, gravity, topology

Any other pairing MUST be rejected during typed-array resolution. constant with any other storage, and unbucketed or time_batch with kind=constant, are malformed typed-array declarations.

4.2 The tag projects the document

dtype and shape appear in both the tag and the document. The tag copy is a readable projection; the digest is authoritative. A reader MUST verify that the projected dtype and shape equal the values in the document the item digest resolves to, and MUST reject the Track on mismatch. A writer MUST NOT emit a tag whose projection disagrees with its document.

5. Identity comparison

Two array modalities are the same type iff their raw wire strings are byte-identical.

Comparison MUST NOT use algo_eq / key_eq. Those compare only the segment after the first namespace, deliberately, so that vortex.* and dreamdb.* algorithm names unify (0002 §5.2). Applied to a modality they would fold distinct namespaces together. Comparison also MUST NOT normalize through the parser first: the parser establishes validity, not identity.

The registry-conflict machinery of 0002 §5.4 applies unchanged — two Manifests registering one array modality with different declarations is a conflict, refused before any content merge.

6. Payload validation

For codec: "raw", the Item's byte length MUST equal product(shape) × sizeof(dtype) exactly. Shorter, longer, or unaligned is corruption and MUST be rejected — not truncated, not padded.

  • product([]) is 1: a scalar occupies one element.
  • A zero extent makes the product 0, and the Item MUST be zero-length. Zero-length dimensions are legal and carry no data.
  • product(shape) × sizeof(dtype) MUST be computed with overflow checking; a declaration whose product overflows 64 bits MUST be rejected at parse time rather than at read time.
  • An all-zero payload is valid. Unlike an embedding under a cosine algorithm (0004 §5), a dense array has no direction requirement and zero is an ordinary value.

7. Codecs

raw is canonical: the elements in layout order, in endianness byte order, with no header, no padding and no alignment beyond the element size. Canonical v1 writers SHOULD emit little and c.

npy exists for interoperability with data already produced by NumPy. When and only when codec is declared npy, the reader parses the NPY header. Then:

  • The header's dtype, shape, byte order and fortran_order MUST match the declared dtype, shape, endianness and layout exactly. A mismatch is corruption and MUST be rejected.
  • The declaration is authoritative; the header is a redundancy check, never an override.
  • A reader MUST NOT inspect magic bytes to decide the codec. Sniffing is exactly the inference this document removes, relocated one layer down.

8. Evolution and merge

An item-type change of any field is a new digest, therefore a new modality, therefore a different Track. It is not a compatible schema evolution and MUST NOT be expressed as one (0017 §2). Migration is the layer mechanism: write the new Track alongside, backfill, and retire the old one.

compatible_with (0017 §2.2) MUST NOT be used to declare two array modalities interchangeable. Two arrays of different dtype, shape or frame are not versions of one field; they are different data.

That prohibition costs arrays nothing, because backfill coverage does not travel through compatible_with for anyone. Coverage of the "write the new Track alongside, backfill, retire the old one" migration above is expressed as a backfill claim (0017 §7) whose basis is the predecessor binding, and compatible_with.coverage is a planner hint that is never consulted for it (0017 §2.2).

Retiring the old Track is subject to 0017 §7.8: while a readable claim names a binding as its basis, that binding stays reachable — moving it from active into lineage is always permitted, dropping it from both is not.

9. Constant Tracks

Per-capture values — a gravity vector, a hand topology, a calibration matrix — are kind=constant, one value for the Track. 0001 §4 defines Constant as one of three Track kinds and 0002 §5.4 gives it an ObjectKind. The Dataset exposes typed-array Constant add/get operations; Python maps them to NumPy arrays and the WASM/TS SDK maps them to JavaScript TypedArrays plus the authoritative shape and semantic metadata. They therefore need neither an invented anchor nor mime="bin".

10. Relationship to 0024

0024 is the identity precedent this document follows. Its corrected wire form carries the complete BLAKE3-256 multihash in one grammatical spec= parameter and rejects truncated, repeated or mismatching identities. Typed arrays apply the same rule to item= and additionally make the entire type document hard and CLOSED; embeddings instead split an identity_basis, which is hashed into the spec_id, from an implementation_record, which is excluded from that hash while still affecting its Manifest's content address (0024 §3).

No compatibility is inferred between an array item= identity and an embedding spec= identity. They hash different document schemas and occupy different modality classes.

11. Post-v1 array families

The family field is the extension axis established by §3. A family has its own CLOSED type document, version, payload magic and validation rules. A dense-v1 implementation presented with either family below MUST reject it as an unknown family; it MUST NOT try to decode the payload as dense v1. This is the compatibility boundary between dense-only consumers and consumers that also support these independently discriminated families.

Both families reuse the built-in array class, the inline item_type registry binding, the complete item= identity, and the legal kind/storage pairs of §§3.1–4.1. Their documents are encoded as deterministic CBOR and hashed exactly as in §3.1. Header and index integers below are unsigned big-endian regardless of the declared element endianness; only values uses that declaration. All byte-length arithmetic is checked u64 arithmetic.

11.1 Ragged arrays — ragged-array version 1

A ragged Item contains a variable number of rows. Each row contains a variable number of fixed-shape cells; inner_shape=[] makes each cell one scalar.

The CLOSED type document contains exactly:

FieldTypeRequired value or meaning
familytextragged-array
versionuint1
dtypetextone dtype from §3
inner_shapearray of uintfixed extents within one cell
endiannesstextlittle or big
layouttextc in version 1
codectextoffsets-values
codec_versionuint1
semantic_type, units, frame, axis_ordertexthard semantic identity fields as in §3

Its canonical modality projection uses shape=r for scalar cells and shape=rx<extents> otherwise (for example, shape=rx2x3). The payload is:

RAG1 | version:u32 | row_count:u64 | cell_count:u64
     | offsets:(row_count + 1) * u64 | values

version is 1. offsets counts cells, not bytes or scalar elements. It MUST start at zero, be nondecreasing, and end at cell_count; equal adjacent values represent an empty row. values has exactly cell_count * product(inner_shape) * sizeof(dtype) bytes. The offsets are inline rather than a sidecar, so an Item is independently decodable and gains no new GC edge.

11.2 Sparse matrices — sparse-csr version 1

Version 1 defines one sparse representation: a rank-2 compressed sparse row matrix. Its CLOSED type document contains exactly:

FieldTypeRequired value or meaning
familytextsparse-csr
versionuint1
dtypetextone dtype from §3
shapearray of two uint[rows, columns]
endiannesstextlittle or big
layouttextc in version 1
index_dtypetextu64
codectextraw
codec_versionuint1
semantic_type, units, frame, axis_ordertexthard semantic identity fields as in §3

The modality projects shape=<rows>x<columns>. The payload is:

CSR1 | version:u32 | rows:u64 | columns:u64 | nnz:u64
     | row_offsets:(rows + 1) * u64
     | column_indices:nnz * u64
     | values:nnz * sizeof(dtype)

version is 1 and the header shape MUST equal the type document. Row offsets MUST start at zero, be nondecreasing, and end at nnz. Within each row, column indices MUST be strictly increasing and each index MUST be less than columns. Therefore duplicate coordinates have no alternate encoding. Explicit zero values are legal and byte-significant; readers MUST NOT delete them or canonicalize them away.

11.3 Still deferred

Compressed payloads, strided/non-contiguous layouts, sparse formats other than CSR, and device-native formats remain undefined. None may reuse either version-1 family while changing its document or payload meaning.

11.4 Storage and range-read properties

The complete encoding above is the Item payload. Event, Continuous and Constant addressing, Fragment packing, chunking, and backfill coverage apply without a family-specific Track or reference. In particular, an inline offset table is metadata within the Item; it is not a content address and adds no closure traversal.

After reading the fixed header, a range-capable reader can locate one ragged row from its two adjacent offsets, or one CSR row from its two row offsets plus the corresponding column and value spans. CSR columns and values occupy separate contiguous regions, so reading one row may require two data ranges. Neither family promises that the returned language value aliases connector memory: byte-order conversion and alignment can require a copy. SDK value types and writer APIs remain future implementation work and MUST preserve the bytes defined here rather than creating a second SDK-specific encoding.

12. Conformance

Vectors are fixture-driven: inputs and expected results a third-party implementation can execute without linking these crates (0009 §3).

Each vector is one JSON document in dreamdb-conformance/vectors/0025/typed-array/, using the exchange envelope of 0009 §4. item_type is the JSON projection of the §3 map (or type_cbor_hex supplies deliberately malformed/non-canonical bytes), modality is the §4 tag, payload_hex is present where payload validation is part of the case, and expected_error names an exact refusal class. Positive identity cases additionally carry expected_type_cbor_hex and expected_item_id. Nothing in a vector depends on a DreamDB-specific API.

Post-v1 format vectors live under dreamdb-conformance/vectors/0025/post-v1-array/ with category post-v1-typed-array. They execute the same §11 codec used by Dataset reads and writes. Positive vectors pin canonical type CBOR, the complete item identity and payload acceptance. Negative vectors name the exact structural refusal.

12.1 Positive requirements

The required cases below may share fixtures where one fixture establishes more than one row; the table is a requirements inventory, not a vector count.

GroupRequired cases
dtype round-tripone per dtype in §3 — 10
rankshape=[] (scalar), 1-D, 2-D, 3-D — 4
endiannesslittle, big — 2
layoutc, f — 2
codecraw, npy — 2
kind/storagethe three legal pairings of §4.1 — 3
zero extenta shape with a 0 extent and a zero-length payload — 1
shared typetwo Tracks, two fields, one item type, no index bound; both resolve (§2) — 1

SDK value vectors carry the complete decoded values for their deliberately small arrays, and every positive vector's modality resolves the TrackKind/ObjectKind. The value matrix catches byte-order and layout errors without requiring large fixtures.

12.2 Refusal and boundary requirements

Each refusal case names the error a conforming reader MUST produce. A vector that merely "fails" does not pass — an implementation that rejects everything would otherwise be conformant. The accepted storage-mapping row is a boundary case, not a refusal.

VectorRequired error
payload one byte shortPayloadLength
payload one byte longPayloadLength
tag dtype ≠ document dtypeProjectionMismatch
tag shape ≠ document shapeProjectionMismatch
NPY header dtype ≠ declarationCodecHeaderMismatch
NPY header fortran_order ≠ layoutCodecHeaderMismatch
NPY magic bytes present, codec: rawPayloadLength — the reader MUST NOT sniff, so it sees only a wrong length (§7)
kind=constant, storage=unbucketedIllegalKindStorage
kind=event, storage=constantIllegalKindStorage
shape product overflows u64ShapeOverflow, at parse time, not at read
item digest truncated to 6 charsMalformedItemDigest
storage=time-batch (hyphen)tag parse error — not a typed-array error (§4.0)
two Manifests, one array modality, different documentsregistry conflict per 0002 §5.4
two fields sharing one array modality that binds an indexSharedIndexedModality (#41)
tag-array-class-unregistered: array.* before the class is addedAmbiguousNamespace — B1
tag-array-missing-kind: array.f32.storage=unbucketed.shape=3.item=…tag rejected — B2, no default
tag-array-missing-storage: array.f32.kind=event.shape=3.item=…tag rejected — B2, no default
tag-storage-hyphen: storage=time-batchtag parse error — B3
storage-mapping-round-trip: each §4.0 rowresolves to the stated ObjectKind — B3

The last two are cross-spec and belong here anyway: they are the cases where typed arrays interact with machinery that already exists, and they are what a reader implemented against this document alone would get wrong.

12.3 What the vectors deliberately do not cover

npy beyond the header fields listed — NumPy's format has variations this document does not adopt, and pinning them would make a partial NPY reader look conformant. The families still deferred by §11.3. And performance: a conforming decode has no required throughput.

13. Implementation blockers

Three findings from checking §4 against the parser. Each is a change an implementation MUST make before any of this document can be honoured, and each has a conformance vector in §12 — they are requirements, not commentary, and are stated here so they cannot be read as background.

B1 — array is unparseable, and not for the reason a reader would guess. ModalityTag::parse rejects array.* with AmbiguousNamespace: the class scan finds neither a known class nor a ≥3-segment reverse-DNS prefix, so it cannot locate the class segment at all. An implementation MUST recognize array as the built-in class declared by 0002 §5.2. Vector: tag-array-class-unregistered.

B2 — kinds must become parameter-derived, which no older class did. Every older built-in class derives TrackKind from the class name and ObjectKind from the class plus a class-specific flag; none reads a declared parameter. §4.1 makes all three kind/storage pairings legal for the single class array, so an implementation MUST extend track_kind() and object_kind() to read the kind and storage parameters, and the typed-array resolver MUST reject a tag that omits either rather than defaulting one. Vectors: tag-array-missing-kind, tag-array-missing-storage, and the three legal pairings of §12.1.

B3 — the storage alphabet and the ObjectKind alphabet are disjoint. storage=time-batch does not parse (hyphens are not legal parameter values) while time-batch is the ObjectKind wire name, and ObjectKind::from_wire("time_batch") returns None. An implementation MUST translate through the §4.0 table and MUST NOT pass a storage value to from_wire. Vectors: tag-storage-hyphen (rejected at parse) and storage-mapping-round-trip (each row of §4.0 resolves).

None of the three is discoverable from this document's prose alone, which is why they are numbered.

14. Verification status

Claims in this document that assert something about the reference implementation, and whether they have been checked against it.

§ClaimStatus
3–7Type document, full identity, registry binding, projections, and raw/NPY validationImplemented and vector-gated — dreamdb-protocol::typed_array; 0025.typed-array.*
4array.* was rejected before implementationHistorical premise verified — AmbiguousNamespace, not the UserDefinedRequiresReverseDns an earlier draft claimed
4Array kinds come from parametersImplemented — ModalityTag::track_kind / object_kind; missing parameters are refused by typed-array resolution
4.0storage=time-batch does not parseVerified — hyphens are not legal parameter values
4.0time_batch is not an ObjectKind wire nameVerified — from_wire("time_batch") is None; the wire name is time-batch
4shape=778x3, shape=0d, base32 item= are legal parameter valuesVerified
4A full tag fits the 256-byte capVerified — MAX_MODALITY_LEN is 256
9Constant TrackKind/ObjectKind exist in the protocol layerVerified
100024 uses a full, grammatical identityVerified — corrected by #192 and gated by 0024.embedding-spec.identity.* vectors
2Several fields may share one index-free modalityVerified — the publish gate refuses sharing only when the modality binds an index
6–7, 12All v1 dtypes, ranks 0–3, both endiannesses/layouts/codecs, all legal kind/storage pairs and zero extentImplemented and vector-gated — 0025.typed-array.sdk-* plus the raw/NPY vectors
9Typed-array Constants are reachable from public SDKsImplemented and behavior-gated — Dataset add/get, Python NumPy round-trip, WASM/TS TypedArray round-trip
12SDKs decode the language-neutral values identicallyImplemented and cross-SDK vector-gated — Python and WASM/TS consume the same non-zero dtype matrix; a WASM-written Dataset containing array Tracks independently opens through the Rust Dataset/CLI path
11Ragged-array v1 and sparse-CSR v1 declarations, identities and payload rejection boundariesImplemented — shared protocol codec, Dataset/SDK append/reopen/point/range reads, physical audit, retained by compaction and GC; 0025.post-v1-array.* plus public Dataset and SDK round trips

Accepted dense v1 does not imply that every implementation supports §11. Compressed, strided and device-native arrays remain deliberately deferred.

14.1 Reference structured-array APIs and resource boundaries

Rust uses ArrayType::{Dense,Ragged,SparseCsr} in ArrayFieldKind. Existing DenseArrayType constructors are accepted by try_new through Into<ArrayType>; item_type() now returns &ArrayType (a Rust source-API change), with as_dense() for dense-specific callers. Dense type bytes and modality identities do not change. Field::Array contains codec bytes; encode_structured builds the §11 payload from offsets, optional columns and numeric bytes and validates it before returning.

Python Schema.add_ragged_array / add_sparse_csr accept component dictionaries with row_offsets, values, and CSR column_indices. Reads return NumPy arrays with declared value dtype and native u64 indices. WASM Schema fields use kind: "array", family, dtype and either innerShape or shape. Writer values use kind: "ragged-array" or "sparse-csr", rowOffsets, values: Uint8Array in declared endianness, and CSR shape / columnIndices. Readers return numeric TypedArrays and BigUint64Array indices; callers need not assemble payload headers.

get_array_item / get_array_item / readArrayItem are Rust/Python/WASM point reads; range/column iterators retain their existing tombstone rules. Physical audit_array_field / auditArrayField reads every referenced Item, including tombstoned Items, using the same content/type checks as reads. Constants use the corresponding Constant APIs. Arrays are indivisible Items: compaction preserves their Tracks without repacking; it does not split a CSR matrix or merge independent ragged values. GC follows the existing Object references; offsets and column indices introduce no references.

The Rust structured view borrows validated payload bytes. SDK transport and component conversion make O(Item bytes) copies, not zero-copy promises; audit holds the Track inventory and at most one payload at a time. Python Arrow projection of these families is explicitly unsupported; typed batch/point APIs remain available. No implicit sparse densification, duplicate coalescing, column sorting or zero elimination occurs.

15. Open Questions

  • OQ-100 (→ this spec): RESOLVED for v1. family is the CLOSED document's discriminant and is exactly dense-array. A later family requires a new specification and is fail-loud to a v1 reader; that specification may retain the array class only if it also defines coherent tag projections, otherwise it uses a new class.

  • OQ-101 (→ this spec): RESOLVED for v1. semantic_type, units, frame and axis_order remain free UTF-8 text and hard identity fields. They compare byte-for-byte with no case folding or vocabulary lookup. Standard vocabularies may be added by a later specification, but v1 never silently equates two spellings.

  • OQ-102 (→ spec/0017): RESOLVED — 0017 §7. Backfilling a new array Track over historical records now has defined coverage semantics. A backfill claim (0002 §7.2.6) names the exact binding being backfilled and an immutable, anchor-unique basis whose Items are the claim's declared domain; a decided-set Object (0002 §7.5.1) carries half-open extents over basis ordinals. A partially backfilled Track is readable, and 0017 §7.5 defines all six read-grid cells — four successful outcomes (Value, Absent, Undecided, OutsideDomain) and a Protocol refusal for the invalid present-but-undecided cell.

    Absent is distinguished from not-yet-written by a positive declaration rather than by the absence of a record: inside a decided range, no record means the backfill decided "no value"; outside it, no record means undecided.

    Ongoing writes are resolved for v1 by 0017 §7.11. An optional field omitted outside every frozen basis reads as OutsideDomain and MUST NOT be mapped to Absent by any API. No implicit claim or moving cohort is created. A caller that needs an explicit absence decision must represent it as data or make the field required; a future opt-in wire form must not reinterpret a v1 omission.