DreamDB

DreamDB Specification — 0002: Content Addressing & Address Grammar

Status: Draft. Builds on 0000-overview.md and 0001-data-model.md. This document fixes: the hash function, the canonical encoding for hashable Objects, the full address grammar, the modality-tag grammar (resolving 0000 OQ-6), and per-entity address derivation. Time-key encoding details are deferred to 0003; spatial-key derivation is deferred to 0004.


1. Purpose

0001 defined what the DreamDB entities are. This document defines what their bytes look like and what their addresses look like — the two prerequisites for content-addressing.

By the end of this document, the following are concrete:

  • The hash function used everywhere in the spec.
  • The canonical byte encoding of every hashable Object (Genesis, Track, Manifest, Bucket/Fragment, etc.).
  • The string form of an address — usable both as a URI and as a backend object key.
  • The modality-tag grammar.
  • How each entity's address is derived from its bytes and its context.

Things still deferred:

  • The time-key encoding (0003) — how a timestamp becomes the time-keyed segment of an address.
  • The spatial-key derivation algorithm (0004) — how a vector becomes a spatial-key bit-string. This document fixes the slot for the spatial key and its encoding, but not how the bits are computed.
  • The per-modality Object byte format (0007) — what's inside a Fragment, a Spatial Bucket, or a Time-bucketed batch.

2. Hash Function: BLAKE3-256

DreamDB uses BLAKE3 (output truncated to 256 bits) as the cryptographic hash everywhere a content hash is needed.

Rationale:

  • Throughput. ~6.8 GB/s single-threaded on modern x86; trivially parallelizable. The 1B-vector ingest path is hash-bound at scale; SHA-256 (~1 GB/s without SHA-NI) becomes the bottleneck.
  • Tree-hashed structure. BLAKE3 is internally a Merkle tree over 1 KiB chunks, which enables efficient subrange verification through the Bao outboard defined in §6.5.4. The root hash alone is not a proof of an arbitrary range: a reader also needs the chaining values on the path from the root to the chunks covering that range. DreamDB therefore stores those values as deterministic derived data and verifies every ranged read before exposing it. Historical Objects without an outboard remain readable only through the whole-Object fallback in §6.5.4. This resolves 0000 OQ-91.
  • 256-bit output. Collision probability ~2⁻¹²⁸ for birthday attacks — adequate forever.
  • Single-spec. No keyed/unkeyed, MAC, or KDF mode confusion at the protocol level. DreamDB always uses unkeyed BLAKE3 for content addressing.

A hash value is therefore always 32 bytes. Wherever this document writes <hash> without qualification, it means "32 bytes of BLAKE3-256 output."

2.1 Multihash prefix

To preserve future hash-function evolution, every hash value that appears as part of a string-form address carries a one-byte algorithm tag:

Algorithm tag (hex)Meaning
0x1eBLAKE3-256 (aligned with IPFS multihash table per 0009 §3.1)
0x12SHA-256 (reserved; not used in v0)
OtherReserved

The tag is prepended to the hash bytes before string encoding (§8). Concretely, every hash in an address is 33 bytes on the wire: 0x1e || <32-byte-hash>.

3. Canonical Encoding: Deterministic CBOR

Every hashable Object in DreamDB (Genesis, Track, Manifest, plus internal sub-structures) is serialized using deterministic CBOR (RFC 8949 §4.2 — "Core Deterministic Encoding Requirements"). The hash is computed over the deterministic-CBOR bytes.

Why CBOR-deterministic:

  • Determinism. Two writers given the same logical Object produce byte-identical bytes (and therefore the same hash). This is what makes content-addressing work across implementations.
  • Compactness. Smaller than JSON; comparable to MessagePack but with an actual deterministic-encoding standard.
  • Tagged types. First-class support for byte strings, integers, maps, and a tag system — no string-escaping gymnastics for binary data like 32-byte hashes or 16-byte nonces.
  • Mature deployments. Used by IPLD-DAG-CBOR, COSE, and CWT.

3.1 Subset rules

To eliminate ambiguity, DreamDB restricts deterministic CBOR slightly beyond RFC 8949:

  • Maps vs. positional arrays — schema-typed DreamDB Objects use one of two encodings, chosen per-schema:
    • Maps with string keys (default for low-volume schemas): self-describing, debugger-friendly. Used for top-level Manifest, Genesis, SpatialIndex Object, and Track Object metadata. Integer-keyed maps are FORBIDDEN.
    • Positional CBOR arrays (mandatory for high-volume schemas — Index Page leaf entries and internal entries per 0007 §7.3 / §7.5): no field names; field order pinned in spec text. Saves ~40 bytes per entry; ~40 GB on a 1B-entry Track index.

3.1.0 Ignorable extensions and critical extensions

§3.1.1 and §3.1.3 each provide a hatch that lets an implementation proceed past an encoding it does not recognise — a trailing element in a positional array, an unknown key in an open map. Neither hatch is unconditional, and this section states the condition both are subject to.

The test. An implementation MAY ignore an unrecognised extension only where ignoring it changes none of the following:

  1. Payload interpretation — which bytes an entry resolves to, and how those bytes are decoded.
  2. Visibility of published data — the set of Items a conforming query returns over data that has already been published.
  3. Reference closure — the set of Objects a collector must mark, including the traversal semantics of references it already reaches.

An extension that changes any of the three is a critical extension. Everything else is an ignorable annotation, and only an ignorable annotation may travel through the hatches of §3.1.1 and §3.1.3.

Condition 3 is written to cover traversal semantics, not merely the location of a reference. An extension that leaves an existing reference field in place while changing what the referenced Object is — from content into an index that must itself be traversed — adds no new reference location and is still critical.

Who this binds. The hatches are worded as reader rules, but the hazard is not confined to readers. This section binds every role that interprets, traverses or reconstructs the structure's semantics:

  • a reader resolving Items;
  • a writer that reads parent or prior state before producing new state;
  • a collector computing a reachable set;
  • an auditor, or a replicator that copies selectively or by reachability, whose output depends on what the structure means.

Byte-exact opaque copying is not bound. A replicator or mirror that transfers Objects verbatim and verifies them against their content addresses (§3.1.2) interprets nothing and ignores no field: it changes none of the three conditions and stays correct without understanding any extension. The obligation attaches only when such a copier stops being opaque — when it selects Objects by meaning, follows a reference closure, or attests that a copy is semantically complete.

It binds specification authors and writers equally: a critical extension MUST NOT be introduced through the ignore hatch of §3.1.1 or §3.1.3, and a specification MUST NOT describe such an extension as forward-compatible.

Consequence for an implementation that does not understand a critical extension. It MUST NOT return a partial, reinterpreted or silently reduced result. It MUST either support the extension as specified, or refuse explicitly before entering the affected semantics — before resolving Items from the Manifest, and before marking or sweeping against it. It need not refuse operations that do not enter those semantics, such as the opaque copying described above.

3.1.0.1 Named historical exceptions

Three critical extensions were specified and published before this section existed, through hatches whose then-stated conditions were insufficient to exclude them — 0014 §2.4 expressly invoked the §3.1.1 hatch, which at the time attached no such condition. They are recorded here as classified under the present test, not as violations of a rule that did not yet exist:

extensionpayloadvisibilityclosure
is_manifest (0014 §2.4)changed—changed
pack_offset (0022 §3.3)changed——
hot_shard (0016 §2.3)—changedchanged

These three are named exceptions to the introduction rule only. The test above applies to every implementation and to every format judgement from this revision onward, including judgements about these three. Being listed here does not make them ignorable. It means:

  • their existing on-disk encodings remain valid formats: they MUST NOT be reinterpreted, deprecated, or made to require rewriting (§3.1.0.2); and
  • an implementation that declares the corresponding capability MUST support it with the semantics its owning specification gives it, while an implementation that does not MUST refuse explicitly before entering the affected semantics, exactly as for any other critical extension.

No further entry may be added to this table. A new critical extension is governed by the introduction rule, not by this list.

3.1.0.2 What this section does not change

This section reinterprets no byte already written. The encodings of is_manifest, pack_offset and hot_shard keep the meanings their owning specifications give them; none is deprecated; no published Object requires rewriting.

Format validity and implementation capability are separate obligations. The chunked (5-tuple) and packed (6-tuple) FragmentEntry shapes and the hot_shard registry key remain valid on the wire, and no implementation may treat them as malformed or as licence to rewrite the Objects carrying them — that is the format obligation, and it is unconditional.

This is not a requirement that every DreamDB implementation implement all three. An implementation without a given capability is conforming so long as it refuses explicitly per §3.1.0.1 instead of degrading. What this revision forbids is silent reinterpretation, not the absence of a capability.

3.1.0.3 Capability floor, not forward compatibility

Because these three shipped without a version boundary, no format mechanism makes an implementation that predates them safe: such an implementation cannot recognise a boundary it was never told about. Their safe use is therefore a property of deployment, not of the format.

A writer MAY continue to emit all three. Doing so declares a controlled deployment capability floor — every reader, writer and collector that will process the affected Manifests supports them — and MUST NOT be described as forward compatibility in an open ecosystem.

The per-role obligations are:

extensionreaderwritercollector
is_manifestresolve the ItemManifest and stitch its chunks, or refusepublish only once the reader and collector floor is metmark the ItemManifest and every chunk it lists, or refuse
pack_offsetfetch [pack_offset .. pack_offset + byte_size), or refusepublish only to deployments that support packed entriesadds no downstream reference; MUST still resolve the entry's format, or refuse
hot_shardmerge HotShard results with cold results, or refuseconfirm the reader and collector floor before publishingmark the HotShard Object, or refuse

Rollback is bound by the same table. Reverting any role to an implementation below the floor requires that the floor no longer be relied upon. For a collector this is assessed over every retained root it will process — tip Manifests, ancestor Manifests and any other Ref — not over the newest Ref alone.

3.1.1 Positional array forward-compatibility (array-length-as-version)

Positional arrays in DreamDB use the array length as an implicit schema version. Future spec revisions (v0.1+) MAY append new fields after the spec'd field count without breaking v0 readers, provided readers obey the discipline:

  • Writers MUST emit at minimum the v0 field count for each schema. Writers MAY append additional fields (per a future spec revision) after the spec'd ones.
  • Readers MUST iterate the array up to their known field count and ignore any trailing fields. They MUST NOT reject an array that has more fields than they expect — this is the forward-compatibility hatch.
  • Readers MUST reject an array shorter than their known field count — this is a malformed write.

This hatch is subject to §3.1.0. A trailing field that changes payload interpretation, the visibility of published data, or reference closure is a critical extension and MUST NOT be introduced this way. is_manifest (0014 §2.4) and pack_offset (0022 §3.3) were introduced through this hatch before §3.1.0 existed; both are named exceptions under §3.1.0.1.

For example, a future schema could append purely diagnostic metadata after all already-assigned positions, if ignoring it leaves all three §3.1.0 properties unchanged. This is not an allocation of a field. A delta that changes how an existing byte size is interpreted would fail that test; positions already used by is_manifest and pack_offset cannot be reassigned as an example extension.

The v0 spec's positional field counts (0007 §7.3 / §7.5) are the minimum for each schema. Future fields go at the end; renumbering the existing positions is a breaking change forbidden within v0.x.

3.1.2 No-re-emit invariant

The forward-compat hatch above creates a subtle hash hazard if SDKs re-encode Objects they didn't author:

  • A v0 SDK fetches a v0.1 Index Page with 5-field leaf entries.
  • The fetched bytes' BLAKE3 matches the address (correct — SDK hashes all bytes received).
  • The SDK decodes entries into its 4-field internal representation, dropping field 5.
  • If the SDK then re-encodes that page (in a derived computation, GC inventory, replicated write, etc.), the re-emitted bytes are 4-field — different bytes, different hash, broken address chain.

To prevent this class of silent corruption:

DreamDB SDKs MUST NOT re-encode Objects they did not originally author. Cached Object bytes are kept verbatim from the backend. Any logical "re-derivation" that produces new bytes also produces a new content hash and is therefore a new Object, addressed at its new hash, never written under the old hash's path.

Operationally:

  • Caches store fetched bytes verbatim (or in a representation that allows byte-identical re-emission); they MAY also keep a decoded representation alongside but MUST NOT prefer it for any output that goes back on the wire.
  • GC walks identify Objects by their backend addresses; they don't decode-and-re-encode for inventory.
  • Federation / mirroring copies bytes verbatim across backends; never re-encodes.
  • Merge operations (per 0008 §5) construct new Manifests with their own hashes; they don't modify existing Manifests.

This applies to all CBOR-encoded DreamDB Objects, not just Index Pages — Manifests, Track Objects, Genesis, SpatialIndex Objects all share the rule. The positional-array forward-compat hatch makes the hazard most visible, but the rule is universal.

3.1.3 Map forward-compatibility (unknown-key tolerance)

DreamDB schema-typed maps (per §3.1's "Maps with string keys" pattern) follow a parallel forward-compatibility rule:

  • Writers MUST emit all spec-required keys for the schema version they target. Writers MAY emit additional, schema-defined OPTIONAL keys, and MAY emit keys defined by a future spec revision.
  • Readers MUST ignore any unknown string key without rejecting the map, subject to §3.1.0. An unknown key is the forward-compatibility hatch for adding new optional ignorable annotations (e.g., space_config was added by a later spec without breaking v0 readers).
  • Readers MUST reject a map missing a spec-required key — this is a malformed write.

This is the map-level analog of the positional-array rule in §3.1.1. The two rules together cover both encoding shapes, and both are bounded by §3.1.0; within that bound, new spec versions extend both array tails and map key-sets safely.

hot_shard and vector_compressor were previously cited above as examples of this hatch. Both are reference-bearing — a collector that ignores either fails to mark a live Object — so neither is a valid example of an ignorable annotation. hot_shard is a named exception under §3.1.0.1. vector_compressor is not classified by this revision: its reader, writer and collector contract has not been audited here, and it requires a separate assessment.

The no-re-emit invariant of §3.1.2 applies identically: a reader observing a map with unknown keys MUST NOT re-encode the map without those keys — that would change the hash. Caches keep bytes verbatim.

Reference-bearing extensions

Unknown-key tolerance is safe only while an ignored key carries no reference. A reader that ignores an unknown key merely loses information; a garbage collector that ignores a key holding an Object address fails to mark a live Object, and the sweep deletes data the Manifest still depends on.

This subsection states condition 3 of §3.1.0 for maps. Conditions 1 and 2 bind equally and independently: a key that carries no reference at all can still be critical by changing payload interpretation or the visibility of published data, and the rules below do not exhaust §3.1.0.

Two rules follow, and they constrain the specification author and the writer — a decoder cannot determine whether an arbitrary 33-byte byte string inside an unknown value is an address or a coincidence, and MUST NOT guess, scan unknown values for address-shaped bytes, or infer references from them:

  • an OPTIONAL key added to an open schema-typed map MUST be non-reference-bearing: it may carry configuration, annotation or policy, never an address that closure traversal must reach.

  • a new reference-bearing location MUST be introduced behind a recognisable version boundary covering the referencing field or map itself, not merely the Object it points at. Reuse of an existing, specification-defined reference field is governed separately below.

    Versioning the target Object is insufficient for a new location: an unknown key added to an open parent map is still invisible to an old collector, which never reaches the target to observe its version. In practice the enclosing structure becomes a discriminated form (§3.1.4). Placing a new reference in an existing field does not qualify merely because a collector already traverses that field.

    Reuse of an existing reference field is not a new reference-bearing location. It is permitted where the field's existing specification already covers this reference shape and its traversal semantics, and it is reuse of a defined field rather than the addition of a new optional reference-bearing key.

    In either case, if the target Object introduces new transitive references or new traversal semantics, the target MUST itself sit behind a version boundary.

    Before publication, every deployed garbage collector MUST satisfy the applicable traversal contract; any collector that does not MUST be stopped or upgraded.

3.1.4 Closed maps (version-discriminated, reference-bearing structures)

Schema-typed maps are open by default: §3.1.3 governs them, and this section does not change that. A spec revision MAY additionally declare a specific structure CLOSED, and only under these conditions:

  • the structure is a version-discriminated container — one carrying an explicit discriminant field (for example "form") whose value names the shape.

    A conforming reader encountering an unrecognised discriminant value MUST reject the whole structure with an explicit error. It MUST NOT interpret it as empty, fall back to a default, or treat it as a Legacy shape. This is a requirement on conforming readers, not a description of how any particular existing implementation behaves.

  • the structure is reference-bearing, so unknown-key tolerance would create the GC hazard described above.

  • the declaration is written into the spec revision that defines the shape.

Nested maps. A CLOSED declaration extends to the schema-typed maps that the same version defines and nests inside the container, even though those inner maps carry no discriminant of their own. Without this, only the outermost map could reject unknown keys and every inner map would remain bound by §3.1.3 — which would leave the hazard exactly where the references live. The declaration MUST enumerate the inner maps it covers.

Reader obligation. A reader encountering an unknown key in a CLOSED map MUST reject the map with an explicit error. It MUST NOT ignore the key, and MUST NOT accept the structure with the key dropped.

Extension. A CLOSED structure is extended by defining a new discriminant value, not by adding keys. Conforming readers that do not recognise the value MUST reject it outright rather than half-understanding it, which is the property that makes closing safe.

Not retroactive. This section closes nothing by itself. Every map defined before the revision that declares it CLOSED remains open under §3.1.3, and no field already required or permitted by an existing specification becomes invalid by virtue of this section existing. A structure is CLOSED only where its own defining revision says so.

3.1.5 General CBOR restrictions

These apply to every CBOR-encoded DreamDB Object, open or CLOSED, and are not scoped to any one sub-section above.

  • Tags — ordinary schema-typed Objects use untagged fields. The local foreign-CBOR convention (§3.2) is not a protocol tag range; external CBOR tags are not interpreted by the protocol.
  • Floats — protocol Objects MUST NOT contain floating-point numbers in fields that affect hashing. (Float canonicalization is a known pitfall.) Vector payloads inside Bucket Objects are not part of this restriction — they live opaquely inside payload byte strings, not as CBOR floats.
  • Indefinite-length items — forbidden everywhere (deterministic CBOR already forbids them; this is a reminder).

3.2 DreamDB CBOR tag

A single locally chosen CBOR tag may be used by agreement when DreamDB byte sequences are embedded in foreign / generic CBOR documents, outside DreamDB's schema-typed Objects. This does not reserve a globally unique or private-use tag.

TagMeaning
dreamdb.tagWraps a DreamDB value only in an agreed application context. That context supplies the inner semantic type; the tag alone does not distinguish, for example, a modality string from a spatial-key string.

The concrete tag value is fixed in 0009 §3.2.

Inside DreamDB's own schema-typed Objects (Genesis, Manifest, Track Object, Index Pages, etc.) this tag is NOT used — the schema's field name (or array position, for positional encoded entries per §3.1) is authoritative about each value's type. The tag exists only for the foreign-CBOR-embedding case.

If your application doesn't embed DreamDB values in foreign CBOR, you'll never encounter this tag.

4. The Address Shape

DreamDB uses one address grammar that is simultaneously the protocol-level identifier of an Item and the backend object key of the Object containing it. The "address IS the path" — the SDK never translates between two addressing systems.

The grammar has the two-part structure introduced in 0000 §5.3, with the bucketing decomposition from §3.1 of 0001:

   <object-address>                        ·  <intra-object-locator>
   ────────────────────────────────────       ──────────────────────
   = <timeline-id>                            optional; present iff
   / <modality-tag>                           the Object is a bucket
   / <spatiotemporal-key>                     containing multiple Items
   / <content-hash>

<object-address> is what is sent to the backend as a get/list key. <intra-object-locator> is consumed by the SDK locally (after the Object is fetched) to extract the specific Item. The two parts are separated by a # (URL fragment-style) when both appear in a string-form address.

4.1 Why this layout

Order of segments matters:

  1. Timeline ID first. Objects are grouped by Timeline, so the Timeline ID is the outermost key. Backends can shard storage on this prefix without coordinating with DreamDB.
  2. Modality second. A query is always against one modality at a time (you do not "search video and audio together"). Modality narrows the address space cheaply.
  3. Spatiotemporal key third. The variable-shape segment that carries the queryable structure (time, spatial, or both — see §6.3).
  4. Content hash last. Disambiguates collisions (per 0000 §5.3) and is the only segment derived from the bytes of the Object itself; it is unknown until the Object is built.

These prefixes organize storage; they do not enumerate a snapshot's query results. Spatial, time-range and combined queries select entries from the chosen Manifest's Tracks and Index Pages, then GET the exact addresses those entries authorize (0005 §5.3.1). LIST can include unpublished, historical or unreachable Objects and cannot establish visibility. In particular, time-only queries on a spatial Track filter its index, and Constants have no time-bucket segment (§6.3).

5. Modality-Tag Grammar (resolves OQ-6)

A modality tag is a structured string that names the type of a Track. It identifies:

  • What the payload bytes mean.
  • Which Track kind the track is (continuous / event / constant).
  • Which Object kind (Fragment / Spatial Bucket / Time-batch / unbucketed) the Track uses.
  • For parameterized modalities, the parameter values (e.g. dimensionality for embedding vectors).

5.1 Grammar

modality-tag    := simple-tag | namespaced-tag
simple-tag      := class "." encoding ( "." param )*
namespaced-tag  := reverse-dns "." class "." encoding ( "." param )*

reverse-dns     := tld-segment "." domain-segment ( "." path-segment )+
class           := lower-alpha-snake
encoding        := lower-alphanumeric-snake
param           := param-key | param-key "=" param-value  ;; e.g. "dim=768"
param-key       := [a-z0-9_]+
param-value     := [A-Za-z0-9_]*
  • Class, encoding, namespace segments, parameter keys and flags use lowercase ASCII bodies ([a-z0-9_]). Parameter values may also contain uppercase ASCII and may be empty. = occurs only once in a key/value parameter.
  • Parsing preserves value bytes and case: field=Name and field=name are not interchangeable. Grammar acceptance does not imply a semantically valid parameter; the owning contract may require a non-empty value, a canonical integer, or another narrower set.
  • Hyphens are forbidden inside segments (so segment boundaries are unambiguously .).
  • Total tag length: max 256 bytes. (Earlier drafts of this spec specified 128; bumped to accommodate realistic reverse-DNS prefixes plus multi-table modality parameters.)

Editorial spelling correction. Parameter names in examples use underscores (spatial_bits, replicate_probes, frag_duration, chunk_size). Their former hyphenated spellings did not satisfy this grammar; correcting an example neither adds a parameter to a writer API nor authorizes rewriting existing tags. Support and semantics are determined by the owning specification. 0010 §4 binds compression through the registry, not a new tag suffix. The value alphabet above retains historical parser acceptance, including that in python-v0.0.7; it does not normalize existing tags or make an empty dim= semantically meaningful.

5.2 Built-in modality classes (v0)

Built-in classes do not require a reverse-DNS prefix. The set is small and fixed in v0:

ClassTrack kindObject kindExample tag
videocontinuousFragmentvideo.h264, video.av1, video.hevc
audiocontinuousFragmentaudio.opus, audio.aac, audio.flac
imagecontinuousFragment (one Object per image)image.jpeg, image.png
embeddingcontinuousSpatial Bucket, unbucketed, OR GraphPageembedding.f32.dim=768.bucketed, embedding.f32.dim=128, embedding.f32.dim=768.graph.r=64
arraykind parameterstorage parameterarray.f32.kind=event.storage=unbucketed.shape=3.item=...
scalareventScalar Bucketscalar.categorical.field=label, scalar.int.field=width
textevent— (no Track provisioned yet; see note)text.utf8
transcripteventTime-bucketed batchtranscript.turn
annotationeventunbucketed (low-volume) OR Time-bucketed batchannotation.json
sceneeventunbucketedscene.boundary
sensoreventTime-bucketed batchsensor.gps, sensor.imu
titleconstantunbucketedtitle.text
authorconstantunbucketedauthor.text, author.json
licenseconstantunbucketedlicense.spdx
sourceconstantunbucketedsource.uri
descriptionconstantunbucketeddescription.text

On the text class. Its indexed forms — text.utf8.bm25, text.utf8.splade — are not Tracks: a TextIndex is an auxiliary Object bound to its source Track through the registry (0015 §3.7), exactly as SpatialIndex and ScalarIndex are. They therefore have no Track kind and no Object kind, and MUST NOT appear in a Manifest's tracks[].

The data form text.utf8 is an Event Track, but no Object kind is pinned here because no implementation provisions such a Track yet. Its Object kind is to be fixed when one is, from the encoding actually chosen — not inferred from the class name. Guessing unbucketed would be actively harmful: 0002 §7.3's inline-index decoder dispatches on Object kind, so a wrong value converts a clean "kind undeterminable" error into a silently wrong decode.

For embedding, the bucketed parameter (presence flag, no =value) selects Spatial Bucket Objects; absence means unbucketed (one Object per vector — only viable for small tracks).

Unlike the fixed class-derived rows above, array derives both kinds from required parameters. Its kind and storage values, their legal pairings, and the explicit mapping from tag spelling to Object-kind wire spelling are defined by 0025 §4–§4.1. Omitting either parameter or using an unlisted pairing is invalid; readers MUST NOT infer a default.

For event classes, the .bucket=<duration> parameter selects Time-bucketed batch storage with the given bucket duration (e.g. transcript.turn.bucket=10s); absence means unbucketed.

5.3 User-defined modalities

Applications that need a modality not in the built-in table use reverse-DNS namespacing:

org.example.diagnostics.heart_rate.f32.dim=1
com.acme.proprietary.scan.bytes
ai.dreamlake.taco.pose.f32.dim=16

The first three segments must form a reverse-DNS path the application controls. This prevents two unrelated applications from accidentally adopting the same modality tag.

Note two easily-missed consequences of the §5.1 grammar:

  • reverse-dns := tld-segment "." domain-segment ( "." path-segment )+ requires at least three segments before the class. ai.dreamlake.pose.f32 is therefore not a valid namespaced tag (only two prefix segments); ai.dreamlake.taco.pose.f32 is.
  • Segment bodies are [a-z0-9_], so hyphens are forbidden inside a class name too: heart_rate, not heart-rate.

Built-in tags are reserved: a user-defined tag MUST NOT use video, audio, image, embedding, array, scalar, text, transcript, annotation, scene, sensor, title, author, license, source, or description as its first segment.

5.4 Track kind & Object kind from the tag alone

A reader that sees a modality tag MUST be able to determine the Track kind and Object kind without consulting any external registry. Implementations achieve this by:

  • Parsing the first segment against the built-in class table (§5.2), or
  • For namespaced tags, requiring that the user-defined class register a Track-kind / Object-kind mapping in the TrackTypeRegistry field of the Manifest's space-config sub-Object (§7).

A Manifest that references a user-defined modality without a corresponding registration is invalid; readers MUST reject it.

5.4.1 Registry encoding: repeated keys (grandfathered)

registry is a CBOR map on the wire. Historical Manifests may repeat its outer modality key: writers once emitted one entry per field keyed by modality, so two fields sharing a modality produced two entries under one key. Current writers no longer produce that shape, but the data exists and must remain readable.

Readers therefore decode registry as an ordered sequence of (key, value) pairs, preserving order and multiplicity rather than collapsing to a unique-key map.

The exception is narrow, and applies only to the outer modality key:

  • a repeated outer key where no entry carries a Track type declaration MUST be tolerated
  • a repeated outer key where any entry carries a Track type declaration MUST be rejected as ambiguous, unless every entry under that key is a legacy index binding
  • duplicate keys inside a registry entry's own map are NOT covered by this exception and MUST be rejected
  • the duplicate-key rejection §7.2.3 declares for the Lineage-v1 maps is unaffected. (§3.1.4 governs unknown keys in CLOSED maps; it states no duplicate-key rule, so there is no general CLOSED-map duplicate rule to invoke)

The second rule is deliberately not "more than one declaration". A single declaration alongside one ordinary entry is already ambiguous, and rejecting only on two-or-more declarations would let that pair through.

Readers MUST NOT resolve a repeated outer key by silently taking the first or the last entry. Two readers once disagreed on exactly that — one took the first match, the other the last — so a field could answer queries from another field's index with no error surfacing. Where a repeated key is ambiguous for the question being asked, the reader MUST reject rather than choose.

Writers SHOULD NOT emit new repeated keys, and MUST NOT rewrite or de-duplicate the entries of a Manifest they did not author (§3.1.2).

SHOULD NOT, not MUST NOT: the writer gate currently accepts two entries under one key when they carry an identical binding, treating them as one binding. Tightening that to a prohibition is a separate change to the gate, with its own test, and is not made here.

This clause is a compatibility exception recording deployed behaviour, not an endorsement. It exists because rejecting these repeated keys outright would make existing Spaces unreadable, which the support policy forbids.

6. Address Components

This section pins down each segment of the address.

6.1 Timeline ID

<timeline-id> := <multihash-of-Genesis-Object>

The Timeline ID is the BLAKE3-256 hash of the deterministic-CBOR encoding of the Timeline Genesis Object (per 0001 §5.1), prefixed with the multihash algorithm tag (§2.1). 33 bytes on the wire; 53 base32 characters in string form (§8).

The Timeline ID is globally unique by the cryptographic argument in 0001 §5.2. Two writers never accidentally share a Timeline; they share one only if they explicitly exchange the Genesis Object.

6.2 Modality tag

The modality tag (§5) appears verbatim in the address as an ASCII segment. Length is bounded at 256 bytes.

6.3 Spatiotemporal key

This is the variable-shape segment. The shape depends on the modality's Object kind:

Object kindSpatiotemporal key shapeNotes
Fragment (media)<time-bucket>Storage-layout hint (see §6.3.1). Placement: floor(t_start / bucket-duration).
Spatial Bucket<spatial-key> or <spatial-key>/<time-bucket>Time-bucket included iff the modality declares spatiotemporal partitioning.
Time-bucketed batch<time-bucket>Storage-layout hint (see §6.3.1). Placement: floor(t_start / bucket-duration).
Unbucketed (Item = Object)<time-anchor>Exact time anchor; encoding per 0003. This is the query primitive (no separate index).
GraphPage listliteral graph-pageVamana-only page Object path; no Item time anchor (0013 §4.2).
VideoItemliteral video-itemLogical video Object path; placement and lookup are carried by the dual outer indexes (design/0014 §4).
Constant(empty)Coverage is "all of time"; modality alone identifies the constant.

<time-bucket> and <time-anchor> encodings are deferred to 0003. <spatial-key> derivation is deferred to 0004; this document fixes only its encoding format:

<spatial-key> := base2( N-bit-string )         ;; literal '0' / '1' chars
                 ;; one character per bit; no padding; no alignment
                 ;; constraint on N. The encoding IS the bit string.

The spatial-key bit length is part of the modality parameters: e.g. embedding.f32.dim=768.bucketed.spatial_bits=18 means 18-bit spatial keys → 18-character keys → up to 2¹⁸ ≈ 262K Spatial Bucket Objects.

6.3.2 Why base2 (and not base32) for spatial keys

Spatial keys carry structured bits whose prefix relationships drive list-prefix queries. Any encoding that aligns to a multi-bit character size (base16 → 4 bits/char, base32 → 5 bits/char) requires padding when the bit length is not a multiple of the alignment, and that padding silently destroys prefix preservation:

Parent (8 bits):     10110011                pad to 10:  1011001100  → "wm"
Child  (10 bits):    10110011 || 10                                  → "wo"
                     ─────────                                          ──
                     extends parent                                     not a prefix of "wm"

The trailing zero pad on the parent contaminates char #2 with bits the child doesn't share, breaking list-prefix-based spatial queries silently.

Base2 sidesteps this entirely: the encoding is literally the bit string, so character-prefix and bit-prefix are the same relationship by construction. A 14-bit prefix query of an 18-bit modality truncates the spatial-key string to 14 chars — no rounding, no encoding gymnastics.

The cost is verbosity: a 20-bit key takes 20 chars instead of 4 base32 chars. Within a DreamDB address that already includes a 53-char Timeline ID and a 53-char content hash, the marginal length is rounding error. Hashes elsewhere in addresses keep base32 (per §8.1) because they are opaque values where conciseness matters; spatial keys keep base2 because they are structured bit strings where prefix semantics matter. Different roles, different encodings.

6.3.1 The time-bucket is a storage-layout hint, not a query primitive

For Fragment- and Time-batch-bearing Tracks, the <time-bucket> segment in an Object's address is purely a storage-layout hint. Its purpose is to give backends a natural prefix on which to shard a million Objects without re-coordinating with DreamDB. It is not the source of truth for time-range queries.

The source of truth is the object_index field on the Track Object (see §7.3), which records the exact time extent of every Object: [(t_start_i, t_end_i, ..., address_i), ...] for Fragments; [(time_bucket_i, batch_address_i), ...] for Time-batches.

Placement rule. A Fragment or Time-batch Object whose contained Items span the time range [t_start, t_end) is placed at the bucket determined by its start:

<time-bucket> := floor( t_start  /  bucket-duration )

A Fragment that crosses a bucket boundary (e.g. covers [59.9s, 60.1s) with 60s buckets) goes into bucket 0. The fact that some of its Items live in clock-time bucket 1 is irrelevant to its placement.

Time-range query rule. A reader answering a time-range query MUST consult the Track Object's object_index — never list-prefix on the <time-bucket> segment alone. Concretely:

  1. Read (and cache per session) the Track Object.
  2. Iterate its object_index, selecting Objects whose [t_start, t_end) overlaps the query range.
  3. Issue ranged-GETs against the matching Objects' addresses.

A query for t = 60.05s against the example Fragment above succeeds because the fragment-index records t_start = 59.9, t_end = 60.1, and interval-containment is checked against those exact values — independent of the storage bucket.

list-prefix(<timeline>/<modality>/<time-bucket>/) is appropriate only as a bootstrap discovery mechanism — when an SDK is enumerating Objects from a backend it has never seen and has no Track Object for. Even in that case, the SDK uses the result as a candidate set and relies on the Track Object (once retrieved) for exact selection.

Recommendation (SHOULD). Writers SHOULD ensure max(item-duration) ≤ bucket-duration for the modalities they emit. Violating this does not break correctness — the object_index handles all cases — but it gradually unbalances storage layout (lots of Fragments straddling boundaries cluster in earlier buckets).

This rule does not apply to Spatial Buckets: every vector has exactly one deterministic spatial-key, so there is no boundary-spanning ambiguity. Spatial-key prefix listing remains a primary query primitive for feature queries.

6.4 Content hash

<content-hash> := <multihash-of-Object-bytes>

The BLAKE3-256 hash of the Object's bytes (whatever the Object's internal format), prefixed with the multihash algorithm tag. 33 bytes on the wire; 53 base32 chars in string form.

For a Genesis Object, the <content-hash> is the Timeline ID — Genesis Objects address themselves by their own hash.

6.5 Intra-object locator

Present only when the Object is a bucket containing multiple Items. All external locators use a single byte-range form — there is exactly one locator syntax in dreamdb:// URIs, regardless of Object kind:

<intra-object-locator> := "bytes:" <start> "-" <end>
                          ;; decimal half-open byte range [start, end)
                          ;; matches HTTP Range: bytes=<start>-<end-1>

The locator is appended to the object-address with a # separator:

<object-address>#bytes:<start>-<end>

For unbucketed Items (Item = Object), the locator is empty and the # is omitted.

6.5.1 Why a single byte-range form

A dreamdb:// URI is self-explanatory at the fetch level: any tool that speaks HTTP Range or an equivalent backend primitive can fetch the bytes without knowing the modality, the Object's internal layout, or anything beyond the URI itself. The modality tag (already in the path) tells the SDK how to decode the bytes once fetched; the byte-range locator tells anyone how to fetch them.

This gives:

  • Universal portability. An SDK that has never seen the modality can still copy, archive, or proxy the Item. It just can't decode the bytes — which is unavoidable without a decoder.
  • Native cacheability. Byte-range URIs map directly to HTTP Range requests, so CDNs and edge proxies cache them transparently. No special-case DreamDB logic needed in intermediate caches.
  • Grammar simplicity. One locator syntax, no per-Object-kind dispatch in URI parsers.

6.5.2 Internal logical references vs external URIs

An SDK MAY reason internally in terms of logical references — idx:1247 for a vector inside a Spatial Bucket, (time-anchor, payload-hash) for an event inside a Time-batch — when it has the Track Object cached and is doing local lookups. What it MUST NOT do is externalize those logical references as dreamdb:// URIs. Anything that crosses the boundary out of the SDK (returned to the application, written into a manifest, shared between SDKs, embedded in another document) MUST be the byte-range form.

6.5.3 Computing the byte range at mint time

Converting a logical reference to a byte range requires knowing the Object's layout. 0007 defines two layout patterns that make this conversion possible without fetching the whole Object:

  • Fixed-size records (the common case). For modalities like embedding.f32.dim=N, every record is a known fixed size; byte_offset(idx) = header_size + idx × record_size. The SDK computes the range from the modality parameters alone.
  • In-Object offset table (fallback for variable-size payloads). The Object begins with a small [(time_anchor_i, byte_offset_i, byte_size_i), ...] table; the SDK fetches just the table (a small ranged GET against the Object's first KB), looks up the record, and emits a URI carrying the resolved bytes: range.

The choice of pattern per modality is fixed in 0007. The Track Object's object_index MAY also carry the per-Item byte ranges inline (avoiding even the small lookup GET) when the modality declares it.

6.5.4 Reader obligation (normative)

A reader resolving an Item whose locator carries a bytes:<start>-<end> range — or whose modality layout implies one per §6.5.3 — MUST fetch exactly the half-open range [start, end) of the object-address bytes and treat those bytes as the Item payload. A reader MUST NOT substitute the whole Object, an Object prefix, or any other range.

This is the load-bearing rule for packed Objects — many Items concatenated under one content-addressed hash (FragmentPacks per 0022, ScalarBucket packs per 0011). Every packed Item shares the same object hash and is distinguished only by its [offset, size). A reader that ignores the offset silently returns the first Item (offset 0) or the entire pack instead of the requested Item — with no error and (since the whole-Object hash still matches) no integrity signal. For fixed-size-record modalities (§6.5.3) the reader MUST derive the same [start, end) it would mint and MUST validate start ≥ header_size, (start − header_size) mod record_size == 0, and end ≤ object_length.

Integrity scope. The whole-Object root alone does not verify which sub-range was returned. A reader MUST authenticate every non-empty ranged read with the standard Bao outboard form below before exposing any byte to the caller or decoder.

For a subject Object whose BLAKE3 multihash is <h>, its deterministic derived outboard lives at:

bao-outboard/<base32(h)>

The outboard bytes are Bao's standard outboard encoding: an unsigned little-endian u64 subject length followed by the preorder sequence of 64-byte parent-node pairs (two 32-byte BLAKE3 chaining values). The outboard is not an ordinary content-addressed DreamDB Object: its path is derived from the subject root, and its bytes are authenticated only when Bao verification reconstructs that root. It MUST NOT be accepted merely because it was fetched from the deterministic path.

For a requested half-open range [a, b) in a subject of length N, where a < b <= N, a reader MUST:

  1. fetch and parse the outboard;
  2. align the data request outward to [floor(a / 1024) * 1024, min(ceil(b / 1024) * 1024, N));
  3. fetch that aligned subject span and verify it with Bao against the 32-byte BLAKE3 digest carried by <h>; and
  4. only after verification succeeds, return exactly [a, b).

The v1 transport MAY fetch the complete outboard. Its size is approximately 6.25% of the subject for large Objects, while the subject transfer remains the aligned range above. Fetching only the proof-node ranges that verification will visit is a compatible transport optimization, not a different proof format or a correctness prerequisite.

Malformed, truncated or root-inconsistent outboards and wrong, truncated or misaligned subject spans are integrity failures. They MUST be refused and MUST NOT trigger the historical fallback: falling back after observing an existing but invalid proof would be a downgrade path. A backend error while reading the outboard also fails the read unless it unambiguously reports Not Found.

Historical compatibility. When and only when the outboard is absent, the reader MUST fetch the complete subject Object, verify its BLAKE3 digest against <h>, and then return exactly [a, b). It MUST NOT return an unproved range. This fallback is functionally conformant but does not claim bounded transfer or bounded memory.

An explicit whole-Object integrity audit or rebuild MAY obtain that complete Object as consecutive raw ranges instead of one response, provided it hashes the entire ordered byte stream against <h> and withholds every decoded result and derived publication until that final comparison succeeds. This exception does not authorize an ordinary query/read API to expose an unproved range, and it does not allow a failed Bao proof to downgrade to whole-Object verification.

A diagnostic structure audit MAY inspect a raw range solely to classify and report malformed published bytes when content integrity is a separate audit dimension. It MUST NOT return those bytes as Item data, use them to answer a query, or publish any derived Object. This narrow observability exception lets an audit report the structural defect that exists instead of replacing it with the earlier fact that its address or proof is also wrong.

Writers MUST publish an outboard for every newly written content-addressed Object family whose specified read path uses byte ranges: Fragment/Time-batch, Spatial Bucket (including partitioned form), Scalar Bucket, Vector-Storage and packed/unbucketed Item Objects. The subject and outboard MUST both be established and pass the same durability barrier before a Manifest or ref may publish a reference to the subject. A failed outboard write may leave an unreferenced subject Object; it MUST NOT advance publication.

The obligation is attached to the Object's specified consumption path, not to the spelling of a reused address family. In particular, a chunk named under a BucketedByTime path but referenced by an ItemManifest and always fetched as a complete content-addressed Object does not require an outboard. If a later protocol reads that chunk by range, new writers of that protocol inherit the obligation and historical chunks use the fallback above.

The outboard carries no independent reachability identity. Garbage collection MUST retain bao-outboard/<h> exactly while subject hash <h> is live, and MAY collect it when the subject is no longer live. This remains true when the same subject bytes occur under more than one content-addressed path: one outboard is shared by their common root.

7. Per-Entity Address Derivation

This section walks through every protocol entity and shows exactly how its address is computed.

7.1 Genesis Object

A Genesis Object is the seed of a Timeline (0001 §5.1). Its CBOR encoding contains:

{
   "origin":         <time-encoding per 0003>,
   "resolution":     <time-encoding per 0003>,
   "horizon":        <optional [t_min, t_max)>,
   "nonce":          h'<16 random bytes>',
   "canonical_name": "<optional UTF-8 string>",
}

Address (= Timeline ID):

<multihash-of-CBOR-bytes>

A Genesis Object has no Timeline ID prefix — it is identity-bearing rather than identity-anchored. Backends store it at the canonical key genesis/<multihash>.

7.2 Manifest

A Manifest enumerates the state of a Space at one moment (0001 §7). Its CBOR encoding contains:

{
   "parents":      [<multihash-of-prev-Manifest>, ...] | [],   ;; empty array for root Manifest
                                                                 ;; one-element array for linear advance
                                                                 ;; multi-element array for merge (see 0008 §2 / §5)
   "timelines":    [<Timeline-ID>, ...],
   "tracks":       <inline-track-list> | <paged-track-list> | <lineage-v1-container> | <lineage-v2-container> | <lineage-v3-container> | <lineage-v4-container>,
   "ts":           <publication time>,
   "writer":       "<opaque writer tag>",
   "registry":     <TrackTypeRegistry sub-object>,    ;; for user-defined modalities (§5.4)
   "space_config": <SpaceConfig sub-object>,           ;; OPTIONAL; absent ⇒ defaults (§7.2.0)
}

The parents field is an array of multihashes, supporting linear advance (one parent), root Manifests (empty array), and merges (multiple parents). DAG semantics are pinned in 0008 §2.

Order is significant and MUST NOT be sorted. For a merge, parents[0] is the trunk and parents[1] is the branch, in the order the merging writer declared them; 0008 §5 defines what distinguishes the two. A Manifest is content-addressed, so the array's order changes its hash: two implementations merging the same pair of parents into the same result MUST agree on the order or they will mint different Manifest addresses for the same merge. Sorting would additionally erase which parent was the trunk, which merge semantics depend on.

A migrated Manifest — one produced by re-ingesting data that previously lived in an older Manifest form — is a root, and carries parents = []. It MUST NOT reference its predecessor: a parent hash addresses a Manifest in the same lineage, and naming a predecessor of a different form would keep that object permanently reachable.

7.2.0 SpaceConfig sub-Object

The OPTIONAL space_config field carries ignorable Space-wide policy such as per-tenant quotas (0018). It is not an authorization trust root: 0012 defines no capability-token issuer keys or key allow-list here. Gateway authorization uses operator-controlled trust and ownership configuration (0018 §3), not issuer keys supplied by a tenant-editable Manifest. This field MUST NOT carry data-plane encryption: encryption changes payload interpretation and reference traversal, so 0019 introduces it through the fail-closed Lineage-v3 discriminant (§7.2.7), not this open map.

It is a CBOR map with the following well-known sub-fields, each independently optional:

"space_config": {
  "quotas":      <Quotas sub-Object>,             ;; per 0018 §2.1
  ...                                              ;; future fields per spec extensions
}

Absent space_config ⇒ no declared quota policy; absence does not authorize a request or establish single-tenant ownership. Sub-fields not present take their documented policy defaults. An unknown legacy federation key remains subject to the open-map rule below; it is not a credential or issuer-trust declaration.

Per the map-extensibility rule in §3.1.1, readers MUST ignore unknown sub-fields of space_config rather than rejecting the Manifest. New spec versions add fields without breaking pre-existing readers.

7.2.1 Inline track list (small Spaces)

For Spaces whose track count is modest (the common case — most Spaces have well under 10,000 tracks), the tracks field is an inline list:

tracks = [
   {
      "address":  <Track-address>,
      "timeline": <Timeline-ID>,
      "modality": "<tag>",
      "kind":     "continuous" | "event" | "constant",
      "role":     "base" | "layer-of:<Track-address>",
      "coverage": [<t_min>, <t_max>),
   },
   ...
]

Writers MUST switch to the paged form (§7.2.2) when the inline list would exceed 1 MiB of CBOR-encoded bytes. Implementations MAY switch sooner.

Each entry also carries an OPTIONAL "field" key — the schema field name this Track belongs to. It is needed when two Tracks share one modality (for example a raw and a preview video.h264 Track in one Space), and implementations have read and written it since before this clause was added; it is documented here to close that gap.

7.2.2 Paged track list (large Spaces)

For very large Spaces, the tracks field uses the same B-tree-of-Index-Pages primitive as Track Objects (§7.3.2):

tracks = {
   "form":         "paged",
   "root":         <multihash-of-root-ManifestIndexPage>,
   "page_count":   <total pages>,
   "total_items":  <total track entries across all leaves>,   ;; named `total_items` on the wire
   "tree_height":  <number of levels>,
   "fanout":       <max children per internal page>,
}

A Manifest Index Page is structurally identical to a Track Index Page (§7.3.2) except that leaves carry track entries (the inline-form record above) and entries are sorted by (timeline_id, modality) rather than by time. Internal pages narrow the search by (timeline_id, modality) ranges.

Manifest Index Pages live at:

manifests/index/<multihash-of-IndexPage-bytes>

Read path for "what tracks does Timeline T have?":

  1. Read Manifest → get root page address.
  2. Recurse on internal pages whose (timeline_id_min, timeline_id_max) covers T.
  3. At leaves, collect track entries with timeline = T.

In practice, paging the Manifest's track list matters only for Spaces with >10K tracks; it is defined here for symmetry with §7.3.2 and to avoid introducing a second mechanism later.

7.2.3 Lineage-v1 tracks container

The inline (§7.2.1) and paged (§7.2.2) forms carry one flat list of Track entries. The Lineage-v1 form additionally distinguishes Tracks that are query-visible from Tracks retained only as provenance, and records the exact parent version each derived Track was built from.

It exists because role: "layer-of:<Track-address>" (§7.2.1) names an exact parent version, and Track addresses are content hashes: once a parent field advances, the exact parent a Layer was built from leaves the query-visible set while remaining necessary for provenance and for garbage-collection reachability.

Lineage-v2 (§7.2.6) extends this form with backfill claims. A writer that emits no claims SHOULD keep using lineage-v1; the two forms are distinct form values and a reader MUST reject the one it does not recognise.

"tracks": {
  "form":    "lineage-v1",
  "active":  [<BindingVersion>, ...],
  "lineage": [<BindingVersion>, ...],
  "edges":   [<LayerEdge>, ...]
}
  • active — the query-visible set. Readers enumerate only this array when resolving Tracks for a query.
  • lineage — provenance-only bindings. Readers MUST NOT enumerate these for queries. Garbage collectors MUST traverse them.
  • edges — the derived-from relation, one edge per derived binding.

active, lineage and edges are all REQUIRED and MAY be empty arrays. Omitting a key is not the same as supplying an empty array, and an omitted key is a malformed write.

CLOSED declaration

Per §3.1.4, this revision declares the following four schema-typed maps CLOSED, and enumerates them as that section requires:

  1. the Lineage-v1 tracks container itself
  2. BindingVersion
  3. BindingVersionRef
  4. LayerEdge

For all four:

  • every key listed below is REQUIRED; there are no OPTIONAL keys
  • a reader encountering an unknown key MUST reject the map with an explicit error, per §3.1.4. It MUST NOT ignore the key, and MUST NOT accept the structure with the key dropped
  • a reader encountering a duplicate key MUST reject the map; it MUST NOT apply a first-wins or last-wins rule
  • extension is by a new form value, never by adding keys

These maps are reference-bearing: address fields name Track Objects that closure traversal must reach. That is why §3.1.3's unknown-key tolerance is not applicable here.

BindingVersionRef

The identity of one binding version. Used inside LayerEdge.

{
  "timeline": <Timeline-ID>,     ;; bstr, 33 bytes, valid multihash
  "field":    "<field name>",    ;; tstr, REQUIRED
  "modality": "<modality tag>",  ;; tstr, MUST parse per §5.1
  "address":  <multihash>        ;; bstr, 33 bytes, valid multihash
}

field is REQUIRED in Lineage-v1, without exception. The logical identity of a binding is the pair (timeline, field); a binding whose field is absent has no logical identity and cannot participate in an edge. A pre-Lineage-v1 entry lacking field may be migrated only where it resolves uniquely; otherwise migration MUST be refused rather than guessed.

A byte string of the correct length is not sufficient: timeline and address MUST each validate as a protocol multihash (§8), since a 33-byte string may still carry an unsupported algorithm tag.

BindingVersion

Identity plus declaration. The element type of active and lineage.

{
  "timeline": <Timeline-ID>,     ;; as above
  "field":    "<field name>",    ;; as above
  "modality": "<modality tag>",  ;; as above
  "address":  <multihash>,       ;; as above
  "kind":     "<track kind>",    ;; tstr: "continuous" | "event" | "constant"
  "coverage": [<t_min>, <t_max>],;; two uints, half-open, t_min < t_max
  "role":     "<role>"           ;; tstr: "base" | "layer-of:<Track-address>"
}
  • kind MUST be a TrackKind wire value. A decoder checks that wire value alone. Manifest structural validation MUST additionally resolve modality against the enclosing Manifest's registry and check that the declared kind equals the resolved TrackKind, for every binding in active and in lineage. Every publish MUST perform that structural validation, so a Manifest whose binding contradicts its own registry is refused at write time rather than at read time.
  • coverage is half-open [t_min, t_max) and MUST satisfy t_min < t_max.
  • role MUST agree with edges biconditionally: a binding with no edge MUST have role == "base", and a binding with an edge MUST carry the layer role defined below. Either half alone is insufficient; both directions are checked.

The layer role is a text rendering of a byte-string address, so the conversion is pinned here rather than left to implementations:

role = "layer-of:" + base32(exact_parent.address)

where base32(...) is the §8.1 encoding — lowercase base32 without padding, exactly 53 characters for a 33-byte multihash.

Nothing else is permitted in the address portion: no whitespace before or after the colon, no = padding, no uppercase, and no other representation that happens to decode to the same multihash. A reader MUST reject a non-canonical rendering even when it decodes to the correct address.

This matters because role is a tstr inside content-addressed bytes. Two SDKs rendering the same address differently would produce different Manifest bytes, and therefore different Manifest hashes, for identical semantic state.

LayerEdge

The element type of edges.

{
  "child":        <BindingVersionRef>,
  "exact_parent": <BindingVersionRef>
}

An edge stores these two keys and nothing else. In particular it does not store the parent's logical binding (timeline, field) separately — that is a projection of exact_parent and would be a second place for the same fact to disagree. Nor does it repeat kind, coverage or role, which are declarations belonging to the binding.

Both child and exact_parent MUST be present, exactly once, in active ∪ lineage.

Canonical ordering

One comparison key:

BindingOrderKey(b) = (b.timeline, b.field, b.modality, b.address)

Each component is compared lexicographically as unsigned bytes — no locale, no case folding, no Unicode normalization. Multihashes compare over their full 33 wire bytes, never over their base32 rendering: base32 ordering and byte ordering are not the same relation. modality is ASCII by §5.1 but is compared as UTF-8 bytes so the rule holds unchanged if that ever widens.

active   MUST be strictly increasing by BindingOrderKey(binding)
lineage  MUST be strictly increasing by BindingOrderKey(binding)
edges    MUST be strictly increasing by (BindingOrderKey(child),
                                         BindingOrderKey(exact_parent))

Strictly increasing, with these consequences:

  • writers MUST sort before encoding
  • readers MUST reject an unsorted array. A reader MUST NOT silently re-sort: the goal is a single portable canonical wire representation for a given state
  • adjacent equal elements are a duplicate and MUST be rejected
Duplicate and structural rejection

A reader MUST reject the container when any of the following holds:

  • two elements of active, or two elements of lineage, share a BindingOrderKey — even if their kind, coverage or role differ
  • two elements of active share a logical binding (timeline, field) while differing in modality or address. At most one version of a logical binding may be declared active at a time. (Two elements sharing the full BindingOrderKey are already rejected by the rule above.) This restriction applies to active only — lineage MAY hold several historical versions of one logical binding, which is the point of retaining it
  • a BindingOrderKey appears in both active and lineage. A binding is query-visible or provenance-only, never both
  • two elements of edges share the full ordering key (child, exact_parent)
  • two edges share a child with different exact_parent values. Each child has at most one exact parent. These are distinct ordering keys, so array ordering alone cannot detect this
  • an edge is a self-edge, child == exact_parent
  • child or exact_parent is absent from, or not unique within, active ∪ lineage
  • the edge relation contains a cycle
  • a lineage binding is not reachable, by following edges, from any active binding that has an edge. Provenance is retained because something live depends on it; an unreachable provenance binding is unexplained
Size limit

The canonical CBOR encoding of the whole Lineage-v1 container MUST NOT exceed 1 MiB (1,048,576 bytes):

size <= 1048576   allowed
size >  1048576   rejected

The count covers the container map and everything inside it — form, active, lineage, edges — and excludes the enclosing Manifest's "tracks" key and all other Manifest fields.

This is the same numeric threshold §7.2.1 applies to the inline track list, but it is not the same measurement: §7.2.1 measures a flat list, this measures the whole container. The two limits are not interchangeable.

Lineage-v1 defines no paged form. A Space whose container exceeds the limit MUST be rejected, not silently paged or truncated; a paged form, if ever needed, is a future form value with its own limit. Writers MUST apply this check before publishing the Manifest, and readers MUST apply it independently so a hand-built Manifest cannot bypass a writer.

7.2.4 Manifest address

A Manifest is stored at the backend key manifests/<multihash-of-CBOR-bytes>. Manifests are not addressed under any Timeline because they reference many Timelines; they live in their own top-level namespace.

The address of the current manifest of a Space is either:

  • The hash itself (hash-addressed Spaces), or
  • A ref pointing at the hash (refs/<ref-name> per §10), for ref-supported backends.

7.2.5 Duplicate keys in Manifest maps

These rules apply to every tracks form — inline (§7.2.1), paged (§7.2.2) and Lineage-v1 (§7.2.3) — and to the Manifest map itself. They are stated here as a peer of those sections, not inside any one of them.

Manifest top-level map. A reader MUST reject a Manifest whose top-level map repeats a key, known or unknown. It MUST NOT process both occurrences, and MUST NOT let a later value overwrite an earlier one.

This differs from registry (§5.4.1) because the situations differ: a writer that produced repeated registry keys is documented in the reference implementation, so that data demonstrably exists. No known reference writer has produced a repeated top-level Manifest key — the key set is fixed and each key emitted once. The rule rests on the absence of any known producer, not on a survey of deployed data; if such data is found, this clause is what must be revisited.

Rejecting also keeps forward compatibility coherent: §3.1.3 retention treats an unknown top-level key as a single value, and a repeated unknown key would have no defined retained form.

Inline track entries (§7.2.1) and the paged tracks map (§7.2.2). Both are open schema maps, so unknown keys are legal in them and their duplicate handling must be defined rather than left to implementations:

  • a repeated known key MUST be rejected. No known reference writer produces such a shape, and silently taking the first or last occurrence is how a reader ends up disagreeing with another reader about the same bytes
  • a repeated unknown key MUST be accepted, and its occurrences retained as an ordered sequence, preserving both their order and their multiplicity. Collapsing them would discard data §3.1.3 requires to survive
  • when re-encoding, occurrences under one key MUST be emitted in their retained order. The sequence is not reordered, sorted or de-duplicated

Occurrences under one key are an ordered sequence, not a set. The protocol does not know what an unknown extension means, so it cannot assume that reordering its occurrences preserves meaning. Two Manifests whose occurrences differ only in order are therefore different, not equivalent — see §5.4.1 for the parallel treatment of registry.

Lineage-v1 (§7.2.3) is CLOSED: every key is known, and any duplicate is rejected by that section's own CLOSED declaration — not by §3.1.4, which addresses unknown keys rather than repeated ones.

7.2.6 Lineage-v2 tracks container

Lineage-v2 is Lineage-v1 plus backfill claims: positive statements about how much of a frozen cohort of Items a field's backfill has decided.

It exists because an optional field a writer omits produces no record, which is byte-identical to a value that has not been backfilled yet. Absence and pending cannot be told apart without a positive declaration, and a reader that guesses picks silently. 0017 §7 defines what a claim means; this section defines its bytes and the container's local rules.

"tracks": {
  "form":            "lineage-v2",
  "active":          [<BindingVersion>, ...],
  "lineage":         [<BindingVersion>, ...],
  "edges":           [<LayerEdge>, ...],
  "backfill_claims": [<BackfillClaim>, ...]
}
Inherited rules

Lineage-v2 incorporates §7.2.3 by reference. BindingVersionRef, BindingVersion and LayerEdge are byte-identical and remain CLOSED; BindingOrderKey, the canonical ordering of active / lineage / edges, and the duplicate-and-structural rejection list all apply unchanged — except for the two differences this section states, D-a (below) and D-b (below).

The rules are referenced rather than restated. Two copies of one list is how they drift.

CLOSED declaration

Per §3.1.4, this revision declares two further schema-typed maps CLOSED:

  1. the Lineage-v2 tracks container itself — exactly form, active, lineage, edges and backfill_claims
  2. BackfillClaim

BindingVersion, BindingVersionRef and LayerEdge are already CLOSED by §7.2.3 and are not redeclared.

For both maps:

  • every key listed is REQUIRED; there are no OPTIONAL keys
  • a reader encountering an unknown key MUST reject the map with an explicit error
  • a reader encountering a duplicate key MUST reject the map; no first-wins or last-wins
  • extension is by a new form value, never by adding keys

active, lineage, edges and backfill_claims are all REQUIRED and MAY be empty arrays. An omitted key is a malformed write.

BackfillClaim
{
  "target":      <BindingVersionRef>,
  "basis":       <BindingVersionRef>,
  "state":       "not-started" | "partial" | "complete",
  "decided_ref": <multihash>
}
KeyMeaning
targetthe exact binding version being backfilled
basisthe immutable binding whose Items the claim undertakes to decide (0017 §7.2)
statea restatement of the decided set, validated against it (0017 §7.6)
decided_refaddress of the decided-set Object (§7.5.1)

state MUST be exactly one of the three strings. Any other value is rejected.

Canonical ordering
backfill_claims  MUST be strictly increasing by BindingOrderKey(claim.target)

Strictly, with the same consequences §7.2.3 states for the other three arrays: writers MUST sort before encoding, readers MUST reject an unsorted array rather than silently re-sorting, and adjacent equal elements are a duplicate and MUST be rejected.

Strict increase also gives the uniqueness rule below its wire form: two claims sharing a target are adjacent equals.

D-b — additional structural rejection

In addition to the §7.2.3 list, a reader MUST reject the container when any of the following holds:

  • two claims share BindingOrderKey(target) — at most one claim per exact target binding version
  • target or basis is absent from, or not unique within, active ∪ lineage. This is the same resolution requirement §7.2.3 imposes on edge endpoints
  • target and basis do not share a timeline
  • target equals basis
  • the claim is not live (below)
Claim liveness

A claim is live iff its target is in active, or its target is reachable from some active binding by following edges — the same reachability §7.2.3 already uses to explain a lineage binding.

A claim on a target that is neither is rejected. This is what stops a claim from bootstrapping an otherwise-dead component into validity.

D-a — a live claim explains its basis

§7.2.3 rejects a container where

a lineage binding is not reachable, by following edges, from any active binding that has an edge.

A basis retained solely to serve a claim may have no LayerEdge from the target: a backfill writes a different field's Track and need not be a derived version of its basis. Under the §7.2.3 rule alone that basis is unexplained and the container is rejected, which would make a conformant claim unrepresentable.

Lineage-v2 therefore extends what counts as an explanation:

A live claim's basis is explained by that claim, and needs no LayerEdge.

The extension runs target → basis, never claim → target. A claim MUST NOT explain its own target, and MUST NOT make an otherwise-unreachable target-plus-claim component valid. Without that restriction a pair of mutually-referencing dead bindings would justify their own retention, and garbage would never become collectable.

Closure traversal

A conformant closure traversal MUST follow, for every live claim, both

  • claim.basis — and therefore that binding's Track address, and
  • claim.decided_ref — the decided-set Object (§7.5.1)

This is stated because it is not implied. A field containing an address is reachable only where a specification says traversal follows it; §7.2.3 states the obligation for the maps it closes for exactly that reason, and Lineage-v2 inherits nothing here.

A claim whose target leaves both active and lineage is removed with it. Only then may that claim's contribution to closure cease — subject to any other live reference to the same Objects.

Size limit

The canonical CBOR encoding of the whole Lineage-v2 container MUST NOT exceed 1 MiB (1,048,576 bytes):

size <= 1048576   allowed
size >  1048576   rejected

This is the same numeric threshold §7.2.3 states for Lineage-v1, declared here directly rather than inherited by implication, and measured the same way: the container map and everything inside it — form, active, lineage, edges, backfill_claims — excluding the enclosing Manifest's "tracks" key and all other Manifest fields.

Claims carry references, never extents, so they contribute a bounded number of bytes each. A Space whose container exceeds the limit MUST be rejected, not silently paged or truncated. Lineage-v2 defines no paged form; a paged form, if ever needed, is a future form value with its own limit. Writers MUST apply this check before publishing, and readers MUST apply it independently so a hand-built Manifest cannot bypass a writer.

7.2.7 Lineage-v3 encrypted tracks container

Lineage-v3 is the fail-closed capability boundary for data-plane encryption (0019). It is Lineage-v2 plus one required encryption policy:

"tracks": {
  "form":            "lineage-v3",
  "active":          [<BindingVersion>, ...],
  "lineage":         [<BindingVersion>, ...],
  "edges":           [<LayerEdge>, ...],
  "backfill_claims": [<BackfillClaim>, ...],
  "encryption":      <EncryptionPolicy>
}

The container is CLOSED and has exactly the six keys shown. Every key is required exactly once; unknown or duplicate keys are malformed. The Lineage-v2 rules in §7.2.6 apply byte-for-byte to active, lineage, edges and backfill_claims, including their CLOSED child maps, ordering, local validation, claim succession, closure traversal and the 1 MiB canonical-container limit. The limit is measured over this six-key container, including encryption and all key slots.

encryption is the CLOSED EncryptionPolicy of 0019 §3.2. Its child KeySlot maps are CLOSED by 0019 §3.3. The policy changes how every address reachable below this cleartext Manifest is interpreted and how reference-bearing Objects are traversed. A reader, parent-reading writer, collector, semantic auditor or reachability-based replicator that does not implement Lineage-v3 MUST reject before entering those semantics, per §3.1.0. It MUST NOT treat this form as Lineage-v2 with one ignored key. Byte-exact opaque copying remains governed by §3.1.0.

An encrypted Manifest MUST use Lineage-v3 even when backfill_claims is empty. Lineage-v3 always means encrypted and has no none mode. Conversely, the inline, paged, Lineage-v1 and Lineage-v2 forms are unencrypted. The form, not envelope magic or an open Manifest key, is the only encryption capability discriminator.

All addresses in active and lineage, plus every content address reached by recursively decoding their Objects, name 0019 envelopes under the one declared policy. Refs, Genesis and the Manifest itself remain cleartext. Closure traversal authenticates and decrypts a reference-bearing envelope before following the references in its plaintext; inability to do so aborts traversal before sweep.

The transition and merge rules are defined in 0019 §2. Changing between an older form and Lineage-v3, or changing encryption mode, suite or domain, is a full Reencode into a migrated root. A successor may update recipient slots within the same mode, suite and domain.

7.3 Track

A Track Object enumerates the Items (or Object-keys, when bucketed) belonging to one Track. Its CBOR encoding contains:

{
   "timeline": <Timeline-ID>,
   "modality": "<tag>",
   "coverage": [<t_min>, <t_max>),
   "object_index": <inline-index> | <paged-index>,
}

For boundary-spanning Object kinds (Fragment, Time-batch), the index records the actual time extent of each Object's contents — which MAY exceed the nominal bucket range when items span a boundary (per §6.3.1). The index is the authoritative source for time-range queries.

7.3.0 Form discrimination (inline vs paged)

The SDK MUST discriminate inline vs. paged form by inspecting the CBOR major type of the object_index field:

  • CBOR major type 4 (array) → inline form (§7.3.1). Iterate entries directly.
  • CBOR major type 5 (map) with field "form": "paged" → paged form (§7.3.2). Descend into the B-tree of Index Pages.
  • CBOR major type 5 (map) with "form": "hot-overlay-v1" → the CLOSED scalar hot overlay of 0016 §2.7. Verify and combine its retained cold Track and typed HotShard references; unsupported consumers MUST refuse it.
  • Any other shape → MUST be rejected as malformed. Surface to caller as a Manifest validation error (0006 §7).

The same rule applies to the Manifest's tracks field (§7.2), extended by the lineage forms:

  • array → inline track list (§7.2.1)
  • map with "form": "paged" → paged track list (§7.2.2)
  • map with "form": "lineage-v1" → Lineage-v1 container (§7.2.3)
  • map with "form": "lineage-v2" → Lineage-v2 container (§7.2.6)
  • map with "form": "lineage-v4" → Optional entity-key container (0026 §2). It inherits Lineage-v2's five CLOSED keys and requires exactly one Timeline and the canonical dreamdb.entity_keys registry root. It does not inherit Lineage-v3 encryption; combining these capabilities requires another specified critical form, not optional keys.
  • map with "form": "lineage-v3" → encrypted Lineage-v3 container (§7.2.7 / 0019)
  • any other shape, including a map whose "form" value is unrecognised → MUST be rejected as malformed

A reader MUST NOT interpret an unrecognised "form" value as empty, fall back to a default, or treat it as one of the known forms.

This discrimination is purely structural — readers do not consult any external registry or metadata to know which form they're parsing. The CBOR shape is self-describing.

7.3.1 Inline form (small tracks)

For tracks whose index fits in a single Object without taxing reads, the index is inline. Each entry is a positional CBOR array (per §3.1's high-volume schema rule), with field order pinned per Track kind below. The same field order is reused by Index Page leaves in the paged form (§7.3.2 / 0007 §7.3), modulo the absolute-vs-delta time encoding (inline form uses absolute u64 anchors; paged-leaf form uses deltas relative to the page's t_min).

Per Track kind:

  • Fragment-bearing Tracks (media): [t_start, t_end, byte_size, fragment_address] per entry. t_end is the exclusive upper bound of the Fragment's time coverage; t_end > t_start. Fragments need not be contiguous (gaps allowed) or non-overlapping (layered Fragments may overlap).
  • Spatial-Bucket Tracks: [spatial_key, t_start, t_end, byte_size, bucket_address, table_id?] per entry — table_id present iff the modality declares tables=L > 1 (per 0004 §6.2). Multiple entries with the same spatial_key are allowed to handle hot-spot spatial keys whose Items have been split across multiple bucket Objects per the splitting rule in 0007.
  • Time-bucketed batch Tracks (events): [t_start, t_end, time_bucket, batch_address] per entry. t_start and t_end are the actual time extent of the items in the batch.
  • Unbucketed Tracks: [item_hash, time_anchor] per entry. item_hash is the 33-byte content multihash; time_anchor is the absolute u64 anchor used to construct the full address in §7.4.4 and to select a time range. A hash alone is not an Item address because the storage path also contains the anchor.
  • GraphPage-list Tracks: [page_hash] per entry, in page-index order. This is the object_kind = "graph-page" shape from 0013 §3.2 and is not an Unbucketed Item Track. It carries no time anchor and MUST NOT be served by the time-range query verb.
  • VideoItem Tracks: the time-ordered index uses the VideoItemEntry shape from design/0014 §4 and is paired with that design's mandatory stable-key index. It is decoded only with a field-bound Track handle resolved to object_kind = "video-item"; the generic Fragment decoder MUST NOT accept it.
  • Geometry Tracks: 0027 §3 defines the dedicated CLOSED GeometryTrack event index, whose anchors identify whole scenes. It requires object_kind = "geometry-item"; the generic Fragment/SpatialBucket decoder MUST refuse it. Its descriptor, page and data references use the global addresses in §7.5.
  • Constant Tracks: a single constant_address (no array wrapper — Constant Tracks have exactly one item).

Forward-compatibility: per 0002 §3.1.1's array-length-as-version rule, future spec revisions MAY append fields after the v0 positions. Readers MUST iterate up to their known field count and ignore trailing fields.

In a paged Unbucketed leaf, the second field is encoded as time_anchor - page.t_min, like the other time-bearing leaf forms. Decoding adds page.t_min back with checked arithmetic. GraphPage-list leaves retain their one-field [page_hash] entries.

Writers MUST switch to the paged form (§7.3.2) when the inline index would exceed 1 MiB of CBOR-encoded bytes. Implementations MAY switch sooner.

7.3.2 Paged form (large tracks: B-tree of immutable Index Pages)

At 1M+ items per track, an inline index becomes a 70+ MB Object — unacceptable for cold-start latency, write churn under live ingest, and warm-cache memory. The paged form replaces the inline list with a small reference to the root of an immutable, content-addressed B-tree of Index Pages.

object_index = {
   "form":         "paged",
   "root":         <multihash-of-root-IndexPage>,
   "page_count":   <total Index Pages in the tree>,
   "total_items":  <total Item-or-Object entries across all leaves>,
   "tree_height":  <number of levels; 1 = root is a leaf>,
   "fanout":       <max children per internal page; default per 0007>,
}

Each Index Page is its own content-addressed Object with the address:

<timeline-id> / <modality-tag> / index / <multihash-of-IndexPage-bytes>

(The literal index segment distinguishes Index Pages from Track Objects, which use track, and from Item Objects, which use <spatiotemporal-key>.)

VideoItem's second outer index is ordered by stable item key and therefore cannot use this time-range page shape. Its CLOSED key-page form and distinct address are defined by design/0014 §4:

<timeline-id> / <modality-tag> / video-item-key-index / <multihash-of-page-bytes>

The CBOR encoding of an Index Page:

IndexPage {
   "type":      "leaf" | "internal",
   "modality":  "<tag>",                    ;; redundant; aids verification
   "t_min":     <time-anchor>,              ;; min t_start in this subtree
   "t_max":     <time-anchor>,              ;; max t_end in this subtree
   "entries":   [<leaf-entry> | <internal-entry>, ...],
                                             ;; sorted by t_start ascending
}

;; leaf-entry shape mirrors the inline form per Object kind:
leaf-entry (Fragment track):
   { "t_start", "t_end", "byte_size", "fragment_address" }
leaf-entry (Time-batch track):
   { "t_start", "t_end", "time_bucket", "batch_address" }
leaf-entry (Spatial-Bucket track):
   { "spatial_key", "t_start", "t_end", "byte_size", "bucket_address",
     "table_id"   ;; optional; present iff modality declares tables=L > 1
   }
   ;; entries sorted by (spatial_key, table_id, t_start)
   ;; multiple entries per spatial_key are expected (hot-spot splitting)
leaf-entry (VideoItem time index):
   `VideoItemEntry` from `design/0014` §4, decoded in VideoItem context

internal-entry:
   { "t_min_in_subtree", "t_max_in_subtree",
     "child_page_address", "child_item_count" }

Read path (find Items overlapping [t_a, t_b]):

  1. Read the Track Object → get root page address. (1 GET)
  2. Recurse: at each internal page, descend into children whose [t_min_in_subtree, t_max_in_subtree] overlaps [t_a, t_b]. (~log_fanout(N) GETs total.)
  3. At leaves, collect matching Item/Object addresses.

For fanout = 256, 1M Fragments → tree height ≈ 3 → ~3 GETs of ~18 KB each ≈ 54 KB total per query (vs. 70 MB inline).

Write path (append a new Item/Object):

  1. Locate the leaf page that should hold the new entry.
  2. Create a new leaf page (old contents + new entry, re-sorted).
  3. Walk up the tree, creating new internal pages whose only change is the updated child pointer.
  4. The new root page address embeds in a new Track Object; old pages remain — they're immutable and may still be referenced by older Manifest versions.

For fanout = 256 and 1M items, append cost ≈ 3 new Pages × 18 KB ≈ 54 KB written per fragment (vs. 70 MB).

Page size and fanout are pinned in 0007 §7.1: default fanout B = 256, target page size 16 KiB, and a 64 KiB cap for pages written by current writers. Readers retain the specified 1 MiB compatibility ceiling for pages emitted by released historical writers. A leaf page that exceeds the target size on append is split; an internal page that exceeds it is also split, propagating the split up if necessary. Splits never modify existing pages — they create new pages and update the tree by copy-on-write up to the root. Untouched historical pages may remain reachable under the compatibility rule in 0007 §7.1.

Spatial-Bucket Tracks special case. The spatial-key already provides log-N navigation via base32 prefix-listing on the backend, so paged indexing for the spatial dimension is unnecessary. However, for very large spatially-bucketed tracks (millions of buckets), the enumeration of spatial-key → bucket-address mappings still benefits from the paged form. The leaf entries are then keyed by spatial_key (lexicographic order) rather than by time, with the same B-tree machinery.

7.3.3 Track Object address

<timeline-id> / <modality-tag> / track / <multihash-of-CBOR-bytes>

The literal track segment distinguishes the Track Object from the Item Objects (under <spatiotemporal-key>/...) and from Index Pages (under index/...) within the same <timeline-id>/<modality-tag>/ prefix.

7.4 Item Objects (per kind)

7.4.1 Fragment (Continuous Signal media)

<timeline-id> / <modality-tag> / <time-bucket> / <multihash-of-Fragment-bytes>

<time-bucket> is floor(t_start / bucket-duration) per §6.3.1 — a storage-layout hint, not a query primitive. A Fragment whose time coverage [t_start, t_end) crosses a bucket boundary is placed by its start; the Track Object's object_index records the exact extent and is consulted for time-range queries.

Writers SHOULD ensure max(fragment-duration) ≤ bucket-duration for balanced storage layout. Internal byte format defined in 0007. Fetched as a whole; sub-frame access is via the byte-range intra-object locator.

7.4.2 Spatial Bucket (Continuous Signal vectors at scale)

<timeline-id> / <modality-tag> / <spatial-key> / <multihash-of-Bucket-bytes>

For spatiotemporally-partitioned modalities:

<timeline-id> / <modality-tag> / <spatial-key> / <time-bucket> / <multihash-of-Bucket-bytes>

Internal format (per 0007): a small header followed by a packed array of vectors with their per-vector time anchors.

7.4.3 Time-bucketed batch (high-volume events)

<timeline-id> / <modality-tag> / <time-bucket> / <multihash-of-batch-bytes>

<time-bucket> is floor(t_start / bucket-duration) per §6.3.1 — a storage-layout hint, not a query primitive. An event whose time anchor is [t_start, t_end) and whose t_end crosses a bucket boundary remains in the batch placed by t_start; the Track Object's object_index records the actual extent of items in the batch.

Writers SHOULD ensure max(event-duration) ≤ bucket-duration for balanced storage layout. Internal format (per 0007): an append-log of (time_anchor, payload) records sorted by time anchor.

7.4.4 Unbucketed Item Object (low-volume events, small embedding tracks)

<timeline-id> / <modality-tag> / <time-anchor> / <multihash-of-Item-bytes>

The Object is the Item; no intra-object locator.

7.4.5 Constant

<timeline-id> / <modality-tag> / <multihash-of-Constant-bytes>

No spatiotemporal key segment; coverage is implicit.

7.4.6 Chunked Item Manifest (per 0014 Path A)

When a blob field's Schema write policy enables chunking (0014 §2.1), each logical Item is stored as N content-addressed chunk Objects plus one ItemManifest Object that enumerates them in byte order:

<timeline-id> / <modality-tag> / items / <multihash-of-ItemManifest-CBOR>

The chunks themselves live at the existing time-bucketed Fragment path (§7.4.1). The Track Object's FragmentEntry for a chunked Item is a 5-tuple with the trailing is_manifest = true (per 0014 §2.4); decoders fetch the map-encoded ItemManifest, then read the listed chunks. Direct entries remain 4-tuples. The flag, not a chunk_size= modality parameter, selects the representation; the reference writer emits no such parameter. Existing Track Object hashes remain unchanged.

7.4.7 Rendition Playlist (per 0014 Path B)

Aligned, already-published VideoItem Objects can be pinned by a RenditionPlaylist snapshot:

<timeline-id> / <modality-tag> / playlist / <multihash-of-RenditionPlaylist-CBOR>

Each rendition directly names a VideoItem modality and content hash. The playlist also commits the common Item key, absolute extent, and relative segment boundaries. A CLOSED registry entry at dreamdb.rendition_playlist.<name> carries exactly v, timeline, modality, and playlist. It is independent of the active field bindings: publishing a playlist does not replace their Tracks. Reachability traversal MUST expand the playlist into each pinned VideoItem's init segment and Fragment or ItemManifest closure (spec/0014 §3).

7.4.8 TextIndex (per 0015 §3)

When a text modality declares an inverted-index algorithm (dreamdb.bm25, dreamdb.bm25-plus, dreamdb.splade-cosine), the index lives at:

<timeline-id> / <modality-tag> / text-index / <multihash-of-TextIndex-CBOR>

Posting-list pages (paged Index Pages per 0002 §7.3.2) live at the parallel slot:

<timeline-id> / <modality-tag> / text-index / posting / <multihash-of-posting-page>

The two-segment text-index / posting form disambiguates the root TextIndex Object from its B-tree of posting-list pages.

7.4.9 Multi-vector retrieval (per 0015 §4)

Multi-vector fields add no Object address slot. Their token vectors live in the existing SpatialBucket and VectorStorage slots under the canonical .bucketed.multivec.maxtokens=M modality. Document identity is recovered from the ordinary record anchor by the composite arithmetic in 0015 §4.2. The previously drafted multi-vec-index/<hash> path and MultiVectorIndex Object were never emitted and are not v1 wire forms.

7.4.10 HotShard (per 0016 §2)

A HotShard buffers recent appends for a Track without forcing a full Manifest publish per Item:

<timeline-id> / <modality-tag> / hot-shard / <multihash-of-HotShard-CBOR>

The Manifest registry's hot_shard field points at the current HotShard. Readers consult both Track + HotShard during query resolution; flush converts HotShard contents into the Track Object via Manifest publish.

7.4.11 GraphPage (per 0013 §4.2)

Graph-indexed modalities (dreamdb.vamana-cosine, dreamdb.fresh-vamana-cosine) store the graph's adjacency-list pages at:

<timeline-id> / <modality-tag> / graph-page / <multihash-of-GraphPage-bytes>

GraphPage Objects are NOT CBOR — they are a custom packed-byte format (per 0013 §4.2: 192-byte header + variable-size node records). Treated like Bucket Objects for storage purposes (raw bytes after a fixed header). The GraphIndex Object's graph_layout.page_node_count field determines which page a node-id lives in.

7.4.12 VideoItem (per design/0014 §5)

Logical videos whose decoder state and Fragment sequence must remain bounded by one stable item identity are stored at:

<timeline-id> / <modality-tag> / video-item / <multihash-of-VideoItem-bytes>

The VideoItem Object is deterministic CBOR. Its CLOSED map, duration and modality cross-checks, init reference, and inline-or-paged relative Fragment index are defined by design/0014 §5. A generic Fragment or Track decoder MUST NOT infer this form from the address or from an init reference; it requires the field-bound object_kind = "video-item" registration.

7.5 Global / cross-Timeline Object Kinds

Several Object Kinds are addressed independently of any single Timeline because they MAY be shared across Timelines (e.g. the same SpatialIndex Object may be referenced by multiple embedding Tracks on multiple Timelines). These live at top-level namespaces:

Path prefixObjectKindDefined inPurpose
manifests/Manifest§7.2Per-Space manifest DAG nodes
spatial-index/SpatialIndex0004 §3.2Partition algorithm params (LSH seed / centroids)
scalar-index/ScalarIndex0011 §4B-tree / bitmap structured-filter index params
vector-compressor/VectorCompressor0010 §3.1PQ / QINCo codebook + weights
graph-index/GraphIndex0013 §3.1Vamana graph metadata + entry point
federation-manifests/FederationManifest0012 §3.1Cross-backend manifest-of-manifests
tenant-usage/TenantUsageBatch0018 §4.1Per-Space rolling resource-usage statistics
refs/Ref / SnapshotTag§10 / 0008 §4.4Ordinary mutable pointers; reserved tags/ roots are create-only
federation-refs/FederationRef0012 §3.3Coordinator-local mutable FederationManifest pointer
tenant-usage-refs/TenantUsageRef0018 §4.3Per-tenant pointer at latest TenantUsageBatch
backfill-decided-set/BackfillDecidedSet§7.5.1Ordinal extents a backfill claim has decided
entity-key-index/EntityKeyIndex0026 §2Typed logical keys, stable identities and revision tokens; not a Track
geometry-item/GeometryItem0027 §3Item descriptor referring to geometry hierarchy and payloads
geometry-page/GeometryPage0027 §3CLOSED cell hierarchy leaf/branch metadata
geometry-data/GeometryData0027 §§1–4Raw packed cell records or complete mesh LOD bytes, with Bao outboard
tombstones/Tombstone head/page0020 §3Flat v1 broadcasts, flat v2 Timeline scope, or paged index objects; follow page roots, parents, and scoped Timeline Genesis

Top-level metadata namespaces use <prefix>/<multihash-of-canonical-CBOR-bytes>; geometry-data/ instead hashes the raw record/LOD bytes defined in 0027. The prefix is part of the path's parser-level discriminator (per §6.3); two ObjectKinds at different prefixes MAY happen to share a hash (statistically improbable, semantically irrelevant — they live in different namespaces).

None of these is a Track. The Track content-shape enumeration (§5.4) is a different axis and MUST NOT be extended to cover them: an auxiliary Object bound to a Track is not a version of that Track. This is the same distinction 0015 §3.7 draws for indexes.

7.5.1 BackfillDecidedSet Object

Addressed at

backfill-decided-set/<multihash-of-canonical-CBOR-bytes>

and referenced by a BackfillClaim.decided_ref (§7.2.6). It holds the extents of the claim's decided set, out of line, so that the number of extents is not bounded by the Lineage-v2 container limit. That is a design choice about unbounded fragmentation, not an assertion that any particular extent set is large: a backfill that swept a cohort without interruption has exactly one range.

The v1 Object is a CLOSED canonical CBOR map with exactly two keys:

{
  "v":      1,                          ;; unsigned integer, exactly 1
  "ranges": [[start, end], ...]         ;; each an array of exactly two u64
}
  • both keys are REQUIRED; there are no OPTIONAL keys
  • an unknown key MUST be rejected with an explicit error
  • a duplicate key MUST be rejected; no first-wins or last-wins
  • v MUST be the unsigned integer 1. A reader that does not recognise the version MUST reject the Object outright rather than interpreting the ranges it can see
  • each element of ranges MUST be an array of exactly two elements, both unsigned integers representable in u64. A one-element or three-element inner array is malformed, not a forward-compatibility hatch — §3.1.1's array-length-as-version discipline does not apply here, and this Object says so explicitly to remove the question
Ordinal extents

start and end are basis ordinals, half-open [start, end) — positions in the basis's ascending-anchor order (0017 §7.3), not time anchors.

Anchors would not be canonical. Over a basis holding anchors {10, 20}, the anchor extents [10, 11) and [10, 20) denote the same set — the first Item decided, the second not — yet encode differently and therefore hash differently. Ordinals are dense over [0, N), so one decided set has exactly one encoding.

Normalisation

A conformant Object satisfies all of:

  • every range has start < end; an empty range is not representable
  • ranges are strictly increasing by start
  • ranges are non-overlapping
  • ranges are coalesced: [a, b) and [b, c) MUST be written as [a, c)
  • an empty ranges array is legal and means nothing is decided

A reader MUST reject an Object violating any of these rather than normalising it on the fly. Silent normalisation would make one decided set reachable at two addresses, which defeats content addressing.

end MUST NOT exceed the basis Item count N; that check needs the basis and belongs to 0017 §7.6, not to decoding these bytes alone.

Cost

v1 is a single Object. Lookup of one ordinal is O(number of ranges), which a contiguous backfill makes O(1). Paging is future work and no mechanism is provided for it here: an implementation that cannot process the Object MUST refuse rather than truncate it.

8. Encoding: base32, paths, separators

8.1 Hash encoding

A multihash (algorithm tag + 32-byte BLAKE3) is 33 bytes. Encoded for use in addresses as lowercase base32 without padding (RFC 4648 §6 with = stripped). 33 bytes × 8 bits / 5 bits per char = ceil(52.8) = 53 characters.

Lowercase base32 is chosen over base64url for:

  • Case-insensitivity. Some object stores normalize case in keys; some filesystems are case-insensitive. Using one case eliminates the issue.
  • DNS-safety. Although DreamDB addresses do not appear in DNS, the same alphabet is friendly to human transcription and many ad-hoc URL handlers.
  • No + / / characters, which conflict with path separators or require URL-escaping.

8.2 Path separator

The / character separates address segments in both the URI form (§9) and the backend object-key form. Backends that use a different native separator (e.g. some flat key-value stores) MUST either accept / literally or be wrapped by an adapter that translates.

8.3 Intra-object separator

The # character separates the object-address from the intra-object locator. This mirrors the URL fragment convention and ensures that backends (which never see anything past #) and SDKs both handle the boundary unambiguously.

9. URI Scheme: dreamdb://

For sharing addresses outside a single backend (e.g. in documentation, manifests of manifests, or chat messages):

dreamdb://[<backend-hint>]/<address>[#<intra-object-locator>]

Where <backend-hint> is an optional hint to the SDK about where to find the bytes — restricted to a single HTTP authority component (host[:port]) and MUST NOT contain /. Backend-specific path components like S3 bucket names are configured via the Connector or expressed via virtual-hosted style addressing (e.g. my-bucket.s3.example.com), NOT embedded in the URI. The hint is purely informational; the address itself is the source of truth, and the SDK MAY consult any backend it knows about.

Examples:

dreamdb:///<timeline-id-base32>/title.text/<hash-base32>
dreamdb://my-bucket.s3.example.com/<timeline-id>/video.h264/<time-bucket>/<hash>#bytes:0-1024
dreamdb:///<timeline-id>/embedding.f32.dim=768.bucketed/<spatial-key>/<hash>#bytes:160-3240

The empty backend-hint form (dreamdb:///...) means "use the default / current backend."

10. Refs Namespace

Ordinary Refs (per 0000 §5.2) are mutable named pointers maintained by ref-conformant backends only. Their backend keys live in a separate top-level namespace, also containing the separately typed tag roots below:

refs/<ref-name>

<ref-name> is a lowercase ASCII path with /-separated segments. Each segment is [a-z0-9_@-], length 1–64. Total ref-name length is bounded at 256 bytes.

@ is permitted so that snapshot refs can carry the <dataset>@<label> shape (0008 §4). At most one @ may appear in a ref name, and it may not begin or end a segment, so the two-part split stays unambiguous.

Examples: refs/main, refs/release/v1, refs/users/alice/scratch.

An ordinary Ref's value is the multihash of the Manifest it points to (33 bytes, base32-encoded for human display, raw bytes for backend storage). Ref updates use conditional writes (CAS / If-Match) per 0000 §5.2.

refs/tags/<name> is reserved for the CLOSED, canonical snapshot-tag-v1 record in 0008 §4.4. It names a Dataset Manifest but is not a raw hash body; ordinary Ref mutators must refuse this subspace. The full tags/<name> obeys the same grammar and length limit. Existing raw Ref bytes are not converted by their name; see the owning rollout/legacy-root rules. Semantic collectors must decode tags or refuse before sweep; opaque copies preserve bytes without claiming to enumerate or certify their closure (§3.1.0).

Cross-backend federated deployments additionally use the federation-refs/ namespace (per 0012 §3.3) for the multi-backend analog. Its CAS semantics are identical to refs/, but v1 has exactly one coordinator-local authoritative copy. Child backends do not race to advance it, and endpoint or credential bindings remain runtime state outside the Federation Manifest.

11. Worked Examples

11.1 A Constant Track

A DreamDB Space records the title "FA Cup Final, 2nd half" for a single timeline.

  • Timeline Genesis: { origin: 2026-05-06T09:00:00Z (ns), resolution: 1ns, horizon: [0,600s), nonce: 0xa3b9..., canonical_name: "match-2026-05-06" }
    • CBOR-encoded → 89 bytes
    • BLAKE3-256 → T_ID = 0x1e || <32 bytes> (33 bytes raw, 53 base32 chars)
  • The Constant Object is the UTF-8 string "FA Cup Final, 2nd half" (22 bytes)
    • BLAKE3-256 → C_HASH = 0x1e || <32 bytes>
  • Address: <T_ID> / title.text / <C_HASH>
    • In string form: xy7g...vqra/title.text/q9ng...uudk (~120 chars)

A query "what is the title?" computes T_ID (one CBOR-and-hash from the Genesis), then issues list-prefix(<T_ID>/title.text/). One backend list, one backend get of the matching key.

11.2 A Spatial Bucket

A 1B-vector embedding track on the same timeline. Modality: embedding.f32.dim=768.bucketed.spatial_bits=18.

  • A vector at t = 152.481 s with payload <3 KB> is hashed by the §6.3 spatial-key derivation (per 0004) to the 18-bit key 101100110011010101 (base2, 18 chars).
  • The vector lands in the bucket whose spatial-key prefix is 101100110011010101. The Bucket Object contains ~3,000 vectors that share this prefix, packed as fixed-size 3080-byte records (8-byte time anchor + 3072-byte f32 vector) after a 160-byte header (per 0007 §6.1).
  • Bucket Object content hash: B_HASH.
  • Address of the Bucket Object: <T_ID> / embedding.f32.dim=768.bucketed.spatial_bits=18 / 101100110011010101 / <B_HASH>
  • The vector's byte offset inside the Bucket: 160 + 1247 × 3080 = 3,840,920. The vector's byte range: [3_840_920, 3_844_000).
  • Address of the individual vector (Item address) externalized as a URI: <T_ID> / embedding.f32.dim=768.bucketed.spatial_bits=18 / 101100110011010101 / <B_HASH>#bytes:3840920-3844000.

A query "find vectors near this one" hashes the query vector to its 18-bit spatial key, issues 1-4 list-prefix(<T_ID>/<modality>/<spatial-key>...) calls (the SDK may truncate to a shorter prefix to widen recall — e.g. truncate to 15 chars for a 15-bit prefix that matches 8× more buckets), fetches the resulting Bucket Objects, and runs exact KNN locally. When returning a result, it emits the byte-range URI shown above; another SDK can fetch that vector's bytes via a single ranged GET, no DreamDB-specific layout knowledge needed.

11.3 A Fragment

A video Fragment for the same timeline. Modality: video.h264. The Fragment covers t = [60s, 62s) and is 1.4 MB.

  • Time bucket: floor(60_000_000_000 / 60_000_000_000) = 1 (assuming 60-second time-buckets, encoding per 0003).
  • Fragment content hash: F_HASH.
  • Address: <T_ID> / video.h264 / 1 / <F_HASH>
  • Address of frame at t = 61.083 s: same as above plus #bytes:412800-415744 (computed from the fragment-index on the Track Object).

Playback: the SDK consults the Track Object's object_index (cached), maps t = 61.083 s to the Fragment's address + byte range, issues one ranged-GET, and feeds the bytes to the decoder.

12. Out of Scope for this Document

  • Time-key encoding — <time-anchor> and <time-bucket> byte/string formats (0003).
  • Spatial-key derivation — how vectors become bit-strings (0004).
  • Per-modality Object internal byte format — Fragment containers, Bucket packed-array layout, Time-batch log format (0007).
  • Conformance test vectors — round-trip CBOR encodings, address-derivation test cases (0009).

13. Open Questions Surfaced by This Document

  • OQ-11 (→ 0009 §3.2): RESOLVED as a local value. dreamdb.tag = 65521 is used only in agreed foreign-CBOR contexts, not ordinary Object schemas. It is not an IANA private-use allocation; §3.2 records that registration limitation.

  • OQ-12 (→ 0009 §3.1): Multihash algorithm tag values. Resolved: 0x1e for BLAKE3-256 (IPFS-aligned).

  • OQ-13 (→ 0007 §6.5): Spatial+time segment order for partitioned modalities. Resolved: spatial-first — <timeline>/<modality>/<spatial-key>/<time-bucket>/<hash>.

  • OQ-14 (→ 0007 §7.1): Default Index Page fanout B and target page size. Resolved: B = 256, target page size 16 KiB, current-writer max page size 64 KiB, historical-reader compatibility ceiling 1 MiB.

  • OQ-15 (→ 0007 §7.2): Inline-vs-paged switch threshold. Resolved: 1 MiB of CBOR-encoded inline bytes; implementations MAY switch sooner.

  • OQ-16 (→ 0007 §6): Per-modality Object layout pattern. Resolved: fixed-size records (default for embedding.f32.dim=N); in-Object offset table (fallback for variable-size payloads, used by Time-batches per §8); reference-mode for multi-table Spatial Buckets (per §6.2).

  • OQ-93 (§5.2, → 0015 §3.7): RESOLVED. Object kind for the text class. Decided, and the five sequenced implementation items — field-qualified index binding, contextual decode, the writer migration off the pseudo TrackEntry, GC reachability, and the Legacy read bridge — have all landed; see INDEX.md for what each one means and where it lives. The retained Legacy read bridge is deliberate, and its removal is separately gated as #121, not unfinished work under this question. The question dissolves rather than being answered: an index is not a Track. Every text.* Track produced by the pre-migration writer was in fact a TextIndex (text.utf8.bm25 / .splade) written into tracks[], which 0015 §3.7 forbids and the current writer no longer does — it adds one registry key and leaves the tracks container untouched: a TextIndex is an auxiliary Object bound to its source Track through the registry, like SpatialIndex and ScalarIndex. Indexes therefore need no Track kind or Object kind at all. The data modality text.utf8 is an Event Track, but no implementation provisions one yet, so its Object kind is deliberately left unpinned rather than guessed — see the note in §5.2.

    Two consequences pinned elsewhere: kind keeps its three values (no discrete — 0015 §3.7's example was corrected), and role keeps its two forms (text-of: / splade-of: / graph-of: do not enter the Track grammar; index relationships live in the registry).


Next: 0003-time-encoding.md — fixes the byte/string format for <time-anchor> and <time-bucket>, and resolves OQ-1 (absolute vs. Genesis-relative origin).