DreamDB Specification — 0006: Protocol Operations
Status: Draft. Builds on
0000-overview.md,0001-data-model.md,0002-content-addressing.md,0003-time-encoding.md,0004-spatial-indexing.md, and0005-backend-interface.md. This document defines the verbs of the DreamDB Protocol — what the SDK does, end to end. It composes the HTTP primitives of0005into multi-step operations, formalizes the per-session cache discipline, and resolves OQ-27 and OQ-28.
1. Purpose
0001–0005 defined the state of a DreamDB Space — what entities exist, how they're encoded, where they live, and what HTTP semantics carry them. This document defines the verbs that operate on that state — the concrete sequences of HTTP requests an SDK generates to ingest, layer, query, and stream DreamDB data.
By the end of this document, the following are concrete:
- The verb taxonomy — the small, well-defined set of operations a conformant SDK exposes to the application.
- For each verb: inputs, outputs, and the canonical HTTP request sequence that implements it.
- The per-session cache discipline — what the SDK MUST cache, MAY cache, and MUST NOT cache, with invalidation rules.
- The concurrency model — what verbs can run in parallel, what guarantees writers and readers have under contention.
- Failure semantics — what happens when a multi-step verb fails partway through, how recovery works, what the backend sees.
What this document does not define:
- The exact application API surface. SDKs in different languages will expose slightly different signatures (Rust async vs. Python sync, etc.). The verb semantics are normative; the surface is implementation-defined.
- The byte format inside Bucket Objects, Fragments, and Time-batches —
0007. - Manifest history walking, branching, and merge semantics —
0008. - Conformance test vectors —
0009.
2. The Verb Taxonomy
The DreamDB Protocol exposes eight verbs in v0:
| Verb | Direction | Primary purpose |
|---|---|---|
Open | Read | Bind to a Space; resolve a Manifest hash or Ref name |
Resolve | Read | Fetch and validate a Manifest by hash; populate cache |
Query | Read | Find Items by time, by feature, or both |
Stream | Read | Fetch a contiguous time range of a media Track |
Get | Read | Fetch a single Item by full DreamDB address |
Append | Write | Add Items to a Track; produce a new Track Object |
Layer | Write | Publish a new derived Track over a parent Track |
Publish | Write | Commit a new Manifest atomically; optionally advance Ref |
Verbs compose: a typical write sequence is Append → Append → Layer → Publish. A typical read is Open → Query → Stream.
2.1 Verbs that are NOT in the taxonomy
These operations exist in lower-level docs but are not first-class verbs:
LIST— only used byOpenfor cold-start bootstrap (0005§5.3.1 Manifest Supremacy). Not a user-facing verb.Update,Modify,Patch,Delete— content layer is immutable; corrections are new layers.Transaction,Lock,Reserve— DreamDB is lock-free; concurrency is handled by content-addressing and ref CAS.Watch— optional SDK composition ofOpen/Resolve(§4.1.2), not a ninth core verb. Backend pushSubscribeand durableStream-changeslogs are not defined.
A conformant SDK exposes the eight listed verbs. Additional convenience methods (e.g., a KNN shortcut wrapping Query with feature input) are implementation-defined and MUST be implementable in terms of the eight.
2.2 Phase-3 / Phase-4 extensions
The following verb is defined in a post-v0 spec draft and is OPTIONAL for v0 SDKs (REQUIRED for SDKs claiming the corresponding Phase-3 / Phase-4 conformance):
| Verb | Direction | Primary purpose | Defined in |
|---|---|---|---|
Reencode | Write | Bulk re-index Items from a source modality to a target modality | 0017 §3 |
Federation (0012) extends Open and Query over an immutable
manifest-of-manifests; it does not add a protocol verb. Copying an Object
closure between backends uses ordinary hash-verified GET/PUT operations and
backend-local authorization.
Reencode is resumable and idempotent — checkpoints published as Layer Manifests every batch_size Items. Crash-and-resume produces the same final state as uninterrupted execution. SDKs implementing it MUST honor the resume_from parameter and verify source-modality non-mutation before continuing.
2.2.1 Query extension: HybridQuery (per 0015 §5)
The Query verb (§4.3) is extended rather than replaced — its query_spec parameter is the structured CBOR sub-Object defined in spec/0015 §5.1, supporting per-modality sub-queries with fusion policy (RRF / linear / max / pareto). A pre-Phase-4 Query invocation with a single dense sub-query is the v0 path unchanged; multi-modality sub-queries activate the spec/0015 query planner.
Resolution of 0015 OQ-66: HybridQuery is not a new verb — it is a richer query_spec payload for the existing Query verb. The eight-verb taxonomy of §2 is unchanged.
A Phase-3 v0.X release SHOULD support 0012's federated read path; a Phase-4
release SHOULD support both Reencode and the extended Query payload shape.
Pre-Phase implementations remain conformant within their phase bounds.
3. The Per-Session Cache
DreamDB's hot-path performance depends on aggressive client-side caching. This section pins down what the SDK caches, when it invalidates, and what guarantees the cache provides.
3.1 What is cached
| Cache slot | Lifetime | Invalidation trigger |
|---|---|---|
| Genesis Object | Per Timeline, until session end | Never (Genesis is immutable; Timeline ID is its hash) |
| Manifest | Per (Space, Manifest hash) pair | Never directly; replaced when SDK Resolves a different hash |
| Track Object | Per Manifest | Invalidated when reading from a different Manifest |
| Index Page | Per Track Object | Invalidated with parent Track Object |
| SpatialIndex Object | Per (modality, hash) pair | Never (content-addressed; immutable) |
| Hyperplane table | Per SpatialIndex Object | Derived; invalidated with SpatialIndex Object |
3.2 Cache identity = content hash
Every cached Object is keyed by its content hash, not by its path or by any "freshness" notion. Two cache entries for the same hash are by definition identical bytes — content-addressing makes consistency trivial.
The SDK MUST NOT cache Objects keyed by anything other than content hash. In particular: caching a Track Object by (timeline, modality) would be a bug, because two different Manifests can have different Track Objects for the same (timeline, modality) pair (e.g., layered corrections).
3.3 What is NOT cached
- list-prefix results. Per
0005§5.3.1 Manifest Supremacy, list-prefix is bootstrap-only. Caching list-prefix results across queries would (a) re-introduce the eventual-consistency window the doctrine eliminates, (b) yield no benefit on the steady-state hot path that doesn't issue list-prefix anyway. - Ref → Manifest mappings, beyond the most recent fetch. Refs are mutable. The SDK MAY cache the latest fetch but MUST NOT treat a stale ref value as authoritative. Refs are re-fetched on
Openinvocations; a long-lived session with a freshResolveafter work-loss is the right pattern. - Bucket Object bytes themselves, beyond what's needed for in-flight queries. A bucket fetched for one query MAY be retained briefly (an LRU window of seconds) but is not part of the protocol-level cache discipline. Aggressive vector caching is an application concern, not a DreamDB concern.
3.4 Cache scope
The cache is per-session, not shared across processes by default. A long-lived application process accumulates a cache that grows with the diversity of Manifests it has consulted; restarting the process clears the cache. SDKs MAY offer a persistent on-disk cache as an extension; this is implementation-defined.
3.5 Cache validation
Because every cached Object is content-addressed, validation is never required. The bytes either match the hash (cache hit, served directly) or they don't (cache corruption — the SDK MUST treat as a bug, evict, and re-fetch). There is no "stale cache" failure mode.
3.6 Cache eviction (memory bounds)
A long-lived SDK process with no eviction policy can accumulate unbounded cache as the application queries diverse Tracks across sessions. At billion-scale across many Tracks the cache could OOM the process: 1000 cached Track Objects × ~5 MB each ≈ 5 GB, plus Index Pages, plus SpatialIndex hyperplane tables.
Conformant SDKs SHOULD provide a bounded cache with a default size cap (recommended: 1 GiB total, configurable). When the cap is exceeded, evict using a standard policy:
- LRU — least-recently-used eviction is the default recommendation.
- Eviction is always safe because cache entries are content-addressed: a re-fetch produces bit-identical bytes against the same hash.
- SDKs MAY use more sophisticated policies (TinyLFU, ARC, segment-LRU) but the default policy MUST be deterministic and bounded.
What MUST stay cached for the duration of an active query:
- The Manifest the query is bound to.
- The Track Object's root Index Page (if paged form).
- Any Index Page currently being traversed.
These are protected from eviction by being held in the query's local stack/working set — outside the cache's eviction policy until the query completes.
The default 1 GiB is a starting point; SDKs running on memory-constrained devices (mobile, edge, embedded) SHOULD lower the default; high-throughput servers MAY raise it. Implementations MUST surface the configured cap to operators (logs, metrics, configuration files).
4. Read Verbs
4.1 Open(space_ref) → Session
Bind to a Space. space_ref is one of:
- A Manifest content hash:
dreamdb://<backend>/<manifest-hash>. - A Ref name (Ref-Conformant backends only):
dreamdb://<backend>/refs/<ref-name>.
HTTP sequence:
Returns a Session handle that wraps:
- The current Manifest (cached).
- The backend URL and auth context.
- A handle to the cache (§3).
The Session is opaque to the application but stable for the duration of subsequent verb calls.
Cold-start bootstrap (no Manifest or Ref known): the application provides a <backend> URL and the SDK issues LIST manifests/ to discover available Manifest hashes. This is the only sanctioned LIST invocation in the read path. The application then chooses which Manifest to bind to (typically the most recent by ts field).
4.1.1 Ref freshness for long-lived sessions
A long-lived Session (a streaming server, desktop application, persistent worker, etc.) bound via Ref MAY drift out of date as another participant Publishes a new Manifest and advances the Ref. v0 is pull-only — there is no push notification — so Sessions interested in seeing new Publishes MUST poll. Two patterns:
Periodic polling (recommended for sessions with an idle steady-state):
The SDK MAY issue a HEAD on the Session's Ref every 30 seconds (default; configurable). If the returned ETag differs from the cached one, the Ref has advanced; the SDK then GET refs/<ref-name> to retrieve the new Manifest hash and Resolves it. Caches built against the old Manifest remain valid for any Object hash they still see in the new Manifest (content-addressing makes cross-Manifest reuse free).
Immediate-before-critical-read polling (recommended for latency-sensitive reads):
If the application requires an absolutely-fresh view for a particular Query, the SDK SHOULD HEAD the Ref before the read if the cached Ref ETag is older than a small threshold (default 5 seconds). One round trip; usually <50 ms. Avoids bound-stale-data correctness bugs in low-latency loops at modest extra cost.
Scope note: Ref freshness polling is unnecessary for short-lived Sessions (e.g., serverless function invocations that Open once and exit). Those are naturally fresh — there is no long-lived cache to drift. Polling is for sustained processes that hold a Session across many queries.
This polling pattern also underlies the optional Watch convenience contract
below. It does not require backend push or change the core Open/Resolve
discipline.
4.1.2 Ref Watch (optional SDK convenience; OQ-34)
Watch observes a named Ref's Manifest pointer, not a log of Publish events. Its inputs are a validated Ref name and an existing authenticated Connector; no wire object, subscription service, credentials format or new protocol verb is introduced. Detached Manifest hashes are immutable and MUST be refused as Watch sources. Constructing a watcher performs no I/O and spawns no worker.
Each poll reads the Ref body and resolves its exact Manifest through the normal
Open verification path. Only after the named bytes pass content-hash and
Manifest decoding checks may the watcher return an observation and update its
in-memory cursor. It does not certify the full reference closure or that every
query capability is supported. ETags MAY optimize transport, but an ETag-only
rewrite of an unchanged Manifest hash is not an observation event.
The first successful poll MUST return the current hash with previous = None.
Later polls return {ref_name, previous, manifest} only when the fetched hash
differs from the last returned hash, or no observation when it is equal.
previous means the previous observation, not the new Manifest's parent.
Rollback and replacement are changes too; no ancestry requirement is imposed.
A → B → C between polls may appear as A → C, and A → B → A may appear unchanged.
Thus Watch provides neither every-publication delivery nor an ordered durable
event stream, exactly-once processing, deletion events or a replay token.
Missing/deleted Refs, transport failures, malformed Ref bodies and invalid or unavailable referenced Manifests MUST be errors, never an unchanged result or a synthetic empty snapshot. Failed or cancelled reads MUST NOT advance the last-observed hash. The same watcher can be retried after repair/reconnection; backoff is caller policy, not a hidden retry loop. A new watcher after process restart starts with a new initial observation. Cached verified immutable Manifests may be reused; a cache hit is not evidence of current remote retention.
Watch MUST NOT refresh any existing Dataset or Session. Consumers wanting a pinned view open the returned Manifest hash, not the Ref name (which could already point elsewhere). Watch creates no GC retention root: retaining old versions for later reads remains the operator's ordinary retention obligation.
The reference Rust SDK exposes RefWatch::poll on native and wasm32 targets,
and Dataset::watch_ref as a convenience constructor. Native next_change
polls immediately and waits a positive, representable caller-specified interval
after each unchanged result; zero/invalid intervals are rejected. Errors return
immediately. Dropping the future cancels its read/wait without a spawned task;
the cursor update and successful return have no intervening await. Connector
timeouts still govern in-flight reads. wasm32 callers supply their event-loop
scheduling around poll; no native timer is emulated. This revision adds no
Python or JavaScript binding and claims no backend push capability.
4.2 Resolve(session, manifest_hash) → Manifest
Fetch a Manifest by hash, validate, cache.
HTTP sequence:
Validation:
- Decode CBOR.
- Verify each
<track-entry>references a known modality (built-in or in the Manifest'sregistry). - Verify each modality's required
spatial_indexreferences (for spatially-bucketed) point at hashes the SDK can fetch. - The fetched bytes' BLAKE3 MUST equal
manifest_hash. Mismatch is a backend error (5xx category from the SDK's perspective).
4.3 Query(session, track_selector, query_spec) → ResultSet
The primary read verb. Three modes determined by query_spec:
Dataset partition-vector queries may tune the selected rerank pool independently
of probe effort, using the query-local option in spec/0010 §5. It is not a
persisted query_spec field or a new wire key. Eligibility precedes pool selection;
Schema exactness is the default; only an explicit mode in 0010 §5 overrides it.
Embedding identity checks remain unchanged. Explicit pool
sizes must be positive and at least the result limit; unsupported query families
reject the option. inherit/approximate/exact modes follow the resolved
OQ-44 contract, including refusal rather than fallback when exactness is unavailable.
| Mode | Inputs | Use case |
|---|---|---|
query_spec.time | Time range [t_start, t_end) | "What events happened between T1 and T2?" |
query_spec.feature | Query vector + recall target ρ | "Find vectors near this query vector" |
Both time + feature | Time range AND query vector | "Find similar vectors recorded in last hour" |
Hot-path HTTP sequence (Track Object cached):
Cold-start sequence (Track Object not yet cached):
The cold-start cost is paid once per Track per session; subsequent queries on the same Track are full hot-path.
4.4 Stream(session, track_selector, time_range) → ByteStream
Fetch a contiguous time range of a media Track as a streaming-decoder-ready byte stream.
HTTP sequence (hot path):
Every ranged GET in step 4 MUST be authenticated before its bytes are emitted,
using the Bao outboard or historical whole-Object fallback in 0002 §6.5.4.
For a FragmentPack entry (0022 §5), the logical Fragment is first resolved at
the pack's fixed time-bucket=0 path and its
[pack_offset, pack_offset + byte_size) range is verified against the pack's
content address. The stream emits that exact logical Fragment, never the whole
pack Object.
The output preserves the stored fragment chain (0007). Decoding also requires
the associated initialization segment: the VideoItem/rendition adapter below
emits it explicitly; callers of the lower-level Fragment Stream supply the
Track's init. The SDK does not transcode or silently concatenate different Items.
Latency budget: First-byte latency ≈ 1 round trip to the first Fragment (~50 ms on commodity HTTP/2). The application can begin decoding the first Fragment while subsequent Fragments are still in flight.
4.4.1 Stream prefetch (look-ahead window)
Object-store tail latency is real. A request for a single Fragment occasionally takes 200–500 ms instead of the typical 50 ms — backend rebalancing, cold-cache misses, transient network hiccups. Without prefetch, every such tail event causes a visible playback stall. With prefetch, tail latencies are masked by the look-ahead buffer.
The SDK SHOULD prefetch ahead of the consumer's current position:
- Default lookahead: 2 Fragments beyond the current consumer position. Aligns with HLS/DASH player conventions; sufficient to mask typical tail latencies (a 2-second Fragment plus 500 ms tail equals ~2.5 s, masked by a 4 s buffer at 2 Fragments × 2 s). The default is 2 (not 3) to minimize wasted prefetch on slow consumers.
- Maximum lookahead: SDK-configurable, default cap 10 Fragments. Larger values waste bandwidth and per-request fees (S3 charges ~$0.0004 per 1000 GETs; a 1000 queries/sec workload with lookahead = 10 costs ~$14K/month in request fees alone). Implementations SHOULD NOT default higher than 10.
- Adaptive sizing (REQUIRED for production): SDKs SHOULD dynamically adjust the lookahead based on observed consumer consumption rate:
- If the consumer is consuming faster than fetches arrive (cache miss), grow lookahead (up to the cap).
- If the consumer is consuming much slower than fetches arrive (cache fill, prefetched Fragments aging out), shrink lookahead toward 1.
- Reset to default = 2 on stream restart or seek. Without adaptive sizing, slow consumers + aggressive default lookahead produce bandwidth and cost waste.
- Cancellation discipline: when the consumer abandons the stream (closes the iterator, errors, or seeks to a different time range), the SDK MUST cancel in-flight prefetch GETs via HTTP/2 stream RST (the
RST_STREAMframe). Leaving in-flight prefetches running wastes bandwidth and (on per-request-priced backends) incurs real charges. - Composition with Stream byte-range fetches: prefetch is a wrapper around the existing
StreamHTTP sequence. The SDK issues N+lookahead concurrent ranged GETs against successive Fragments (HTTP/2 multiplexed); the consumer-side iterator yields bytes in order as they arrive.
Cost-aware default: the spec's "default 2, cap 10, adaptive" stance is calibrated to be safe-by-default on per-request-priced backends. Implementations targeting unmetered backends (self-hosted MinIO, fully-paid CDNs) MAY raise the default lookahead, but MUST honor the cancellation discipline regardless.
This is a performance hint, not a protocol-level guarantee. Conformant SDKs MUST function correctly with lookahead = 0 (synchronous fetch); but production-grade implementations SHOULD prefetch to mask backend tail latency.
4.4.2 Explicit VideoItem / rendition selection (OQ-61)
An SDK media Stream selector MUST distinguish an ordinary field-local VideoItem
key from a named RenditionPlaylist; matching strings in the two namespaces are
not a basis for guessing. For a playlist, an omitted label selects the committed
default_rendition; an explicit label selects that exact label. Unknown labels
MUST fail, never fall back. Missing/invalid default metadata is a malformed
playlist, not permission to pick its first entry. The existing CLOSED playlist
decoder and read-time alignment checks in 0014 §3.2–§3.3 apply unchanged.
Ranges are half-open and relative to the selected Item. Empty/reversed or out-of-duration ranges are refused. Playlist ranges MUST use the shared Fragment boundaries, so a switch is explicit. A plain VideoItem may select complete overlapping Fragments. A missing plain Item key yields an empty stream; a missing playlist or invalid selector yields an error.
The reference Rust Dataset::stream_media(MediaStreamSelector, range) returns
an ordered stream of MediaStreamPart: exactly one Init carrying Item
metadata, the requested range and verified initialization bytes, followed by
complete verified Fragment results. Metadata is validated before init is
emitted. Each later Fragment is fetched on demand; a read/integrity failure
terminates the stream without skipping the missing media. Previously emitted
bytes cannot be retracted, so this is not an atomic whole-range read.
Selection is bound to the Dataset's Manifest and the playlist's exact historical bindings. Another writer changing its Ref, current field binding or playlist default MUST NOT change an existing reader's choice. A caller explicitly opens a newer snapshot to adopt those changes. All selectors reuse the existing VideoItem/rendition selection and verification logic; no codec or playlist wire format is added, and no adaptive bitrate decision is made by this API.
This initial adapter uses no payload prefetch or background tasks. Dropping its
stream cancels the current future and starts no later fetch. It may materialize
metadata and a complete logical Fragment (including reconstructed chunks), so
it claims neither a total RSS bound nor one-round-trip cold-start latency.
The lower-level protocol Stream for legacy continuous Fragment Tracks remains
unchanged; the Dataset iter_stream table-scan API is not this media verb.
The adapter is available to native/wasm Rust consumers; this revision does not
claim new Python/JavaScript bindings or HTTP/2-specific cancellation behavior.
4.5 Get(session, item_address) → Bytes
Fetch a single Item by its full DreamDB address (per 0002 §6.5).
HTTP sequence:
Returns the bytes directly — opaque to DreamDB, decodable by whatever interprets the modality (the application).
This is the simplest verb: one address in, one byte sequence out, one HTTP request. Used when the application has already located an Item via Query and wants its full payload.
5. Write Verbs
5.1 Append(session, track_selector, items) → AppendResult
Add items to an existing Track or create a new Track. items is a sequence of (time_anchor, payload) pairs (or single Constant for Constant Tracks).
Zero-Item handling: per 0001 §4.5, Append with zero Items SHOULD be a no-op — the SDK returns success without producing a new Track Object or Manifest. The application calling Append with empty input is expected to skip the subsequent Publish call.
Constant Track Append: per 0001 §4.5 the Track must have exactly one Constant. Append to a Constant Track MUST be called with exactly one Item; zero or two+ items is a programming error and the SDK MUST reject the call.
HTTP sequence (Continuous Signal — Spatial Bucket case):
HTTP sequence (Continuous Signal — Fragment / media):
HTTP sequence (Discrete Event — bucketed):
Standard write ordering (per 0005 §5.3.1): leaf Objects first (Buckets, Fragments, Batches), then Index Pages, then Track Objects. This ensures Manifest Supremacy: a reader resolving a future Manifest that references the new Track will find all dependencies live.
5.2 Append — atomicity and partial failure
Append is not atomic at the protocol level. Each PUT in the sequence is independently atomic; the sequence as a whole is not. Concrete implications:
- If
Appendfails at stepN, steps1..N-1have left content-addressed Objects on the backend that no Manifest yet references. These are orphan Objects. They cost storage but corrupt nothing. - The application MAY retry: re-running
Appendwith the same items produces identical content hashes (deterministic CBOR, deterministic spatial keys). PUTs for already-present Objects are idempotent (412 → success). Effectively, retries pick up where the failure left off. - Operator-level GC (per
0005§3.6) periodically reclaims orphans whose Manifests were never published.
No multi-PUT atomicity is required from the backend. The protocol's correctness comes from the leaf-first ordering plus content-addressing — not from transactional semantics.
5.3 Layer(session, parent_track_selector, derived_track_spec) → LayerResult
Publish a new Track that derives from an existing parent Track on the same Timeline.
HTTP sequence:
The new Track is structurally a regular Track; the layer relationship is a Manifest concern, declared in the Manifest's tracks entry (per 0001 §6 and 0002 §7.2).
5.4 Publish(session, manifest_spec) → PublishResult
Atomically commit a new Manifest. This is the only verb that changes the visible state of the Space.
HTTP sequence:
Concurrency under Ref CAS: two concurrent writers attempting to advance the same Ref will see one succeed; the loser receives 412 and either retries or rebuilds the Manifest with the winner's Manifest as the new parent. This is the optimistic-concurrency story from 0000 §5.2 made concrete.
Hash-addressed Spaces (no Ref): Publish simply PUTs the new Manifest. Discovering it requires out-of-band manifest distribution (the writer hands the hash to readers via some channel — Slack, email, another system). Two concurrent writers in this mode produce two diverging Manifests; reconciliation is an application concern.
5.5 Multi-Object writer transactions in summary
A typical writer transaction:
Phase 1 PUTs leaf Objects, Index Pages, and Track Objects — all content-addressed and idempotent. If the process crashes between Phase 1 and Phase 2, the Phase 1 Objects are orphaned (no Manifest references them), but the Space's visible state is unchanged. A retry of the entire transaction reproduces identical bytes (deterministic encoding) and Phase 1 PUTs become no-ops; only Phase 2 makes progress.
This is the lock-free collaborative pattern from 0000 §5.2 in practice: any number of concurrent writers can stage Phase 1 work in parallel without coordination, and the only contention point is the Ref CAS in Phase 2.
6. Concurrency Model
6.1 Reader-reader concurrency
Trivially safe. Readers share no mutable state; backend GETs are idempotent. Multiple SDK sessions on the same Space can issue parallel Query and Stream calls without coordination.
6.2 Reader-writer concurrency
Safe by Manifest Supremacy. A reader bound to Manifest M_n is unaffected by a writer publishing M_{n+1} — the reader continues to resolve via M_n's object_index. The reader sees a consistent snapshot of the Space at M_n until the application chooses to Resolve a newer Manifest.
A reader that wants to "see new writes" calls Resolve(latest_hash) (or fetches the Ref again). Until that call, the reader is operating on a stable snapshot.
6.3 Writer-writer concurrency
The lock-free pattern. Two writers W_a and W_b:
- Both stage Phase 1 in parallel. Their leaf Objects, Index Pages, and Track Objects are content-addressed and PUT independently. PUTs are idempotent; even if both writers happen to compute the same content (rare but possible — same modality, same Items at same time anchors), they produce identical bytes and the second PUT is a no-op.
- Both attempt Phase 2. If both target the same Ref, one wins the CAS, one loses (412). The loser:
- Re-fetches the Ref → new winner Manifest hash.
- Resolves the winner's Manifest.
- Rebuilds its own Manifest with the winner as the new parent, preserving its staged tracks.
- Re-attempts the CAS.
The resolution preserves both writers' work in a serialized history. The losing writer's Phase 1 Objects are NOT thrown away — they're still content-addressed Objects on the backend, and the rebuilt Manifest references them.
For hash-addressed Spaces (no Ref), there's no central coordination point; concurrent writers produce diverging Manifests and any reconciliation must happen out-of-band.
6.4 Adaptive recall widening (resolves OQ-28)
Per 0004 §6.5 (the combined recall-widening procedure; renumbered from §6.4 in 2026-05 when read-time multi-probe was promoted to first-class Lever 4), the SDK MAY iteratively widen the spatial-key prefix-truncation depth M if the initial result set is too small. v0 makes this implementation-defined with two guidelines:
- The SDK SHOULD start with a
Mthat achieves the requested recall ρ for a defaultθ_max = 30°(or the modality's declared default, if any). - If the result set has fewer than
k(the user's requested top-k) candidates, the SDK MAY decrementMand re-issue prefix queries against the cachedobject_index. (No additional list-prefix round trip — the index is already cached.) - The SDK MUST stop widening at
M = 0(full-track scan). ReachingM = 0indicates either a query vector with no near neighbors in the Track (legitimate) or a misconfigured modality (not DreamDB's concern).
Adaptive widening composes naturally with multi-table — each table widens independently.
6.5 Speculative preloading (resolves OQ-28)
Per 0004 §7.3 and 0005 §3.5, the SDK MAY issue speculative GETs for "most likely" Bucket / Fragment Objects while a list-prefix call (cold-start path) is still in flight. v0 makes this implementation-defined:
- The SDK MUST NOT issue speculative GETs for paths it cannot prove exist in the cached
object_index(which would generate spurious 404s, wasting backend calls and possibly costing money under per-request pricing). - The SDK MAY issue speculative GETs for paths derived from a partial result of a paginated list-prefix (i.e., start fetching the first page's hits while later pages are still loading).
- Cancellation: if a speculatively-fetched Object turns out to be unneeded, the SDK SHOULD cancel the in-flight HTTP/2 stream rather than waste bandwidth.
7. Failure Semantics
7.1 Per-verb failure modes
| Verb | Failure modes | Recovery |
|---|---|---|
Open | Backend unreachable, manifest 404, ref 404, hash mismatch | Retry; if persistent, surface to application |
Resolve | Manifest 404, hash mismatch, CBOR malformed | Treat hash mismatch as critical (corrupt data); 404 likely a typo or removed Manifest |
Query | Track Object missing, bucket 404, decode error | See §7.4 — surface as ObjectNotFound; do NOT silently treat as zero results |
Stream | Fragment 404 mid-stream | Surface as ObjectNotFound; do not silently skip Fragments |
Get | 404 | Surface as ObjectNotFound |
Append | PUT failure mid-sequence | Retry from the failed step; idempotent PUTs make retry safe |
Layer | Same as Append | Same |
Publish | Manifest PUT 5xx | Retry. Don't advance Ref until Manifest PUT succeeds. |
| Ref CAS 412 | Per §6.3 — rebuild Manifest with new parent, retry CAS |
7.2 Crash recovery
A crash mid-Append leaves orphan Objects but no inconsistency. A crash mid-Publish (Manifest PUT succeeded but Ref CAS hadn't completed) leaves the new Manifest content-addressed on the backend; the next Open call won't find it via Ref but can be told its hash directly. Application-level recovery (write-ahead logging the staged Manifest hash before attempting Ref CAS) is implementation-defined.
7.4 ObjectNotFound — error class for missing Objects
When an SDK fetches an Object referenced by a Manifest and receives 404 Not Found, the SDK MUST surface this as a distinct error class — ObjectNotFound — to the application. It MUST NOT silently treat the missing data as zero results, hide the error, or retry indefinitely.
The error MUST carry:
- The full DreamDB address of the missing Object (the address that returned 404).
- The Manifest hash from which this Object was reachable (the path that led the SDK to this Object).
- The Object kind (Bucket / Fragment / Time-batch / Track Object / Index Page / etc.) inferred from the address path.
7.4.1 When ObjectNotFound legitimately occurs
The ObjectNotFound condition is abnormal but bounded — it indicates that some operator-level event has caused an Object to disappear from the backend while a Manifest still references it. Legitimate causes:
- GC race (per §7.3.2.1): a long-running write transaction's Phase-1 Object was GC'd before its Manifest was published.
- Operator-deleted Ref: a Ref was explicitly deleted, GC ran, and an SDK with a cached Manifest hash from that retired Ref tries to query.
- Cross-backend federation gap: a Manifest was federated from one backend to another, but the transitive Object closure wasn't completely copied yet.
- Corrupt operator action: an operator accidentally deleted Objects without GC's reachability check.
In all these cases, the data is genuinely missing — there is no recovery the SDK alone can perform.
7.4.2 Application-level handling
Applications encountering ObjectNotFound have three reasonable strategies:
- Propagate: surface the error to the user / caller. The query failed; show a meaningful error message naming the missing data.
- Fall back: the application has a different Manifest (an older snapshot, a different Ref) that still has all its Objects intact; retry against that.
- Recover (operator action): identify the missing Object's logical content from external state (an upstream pipeline's output, a backup), re-PUT it. Same content → same hash → restored. The Manifest becomes resolvable again.
DreamDB's role ends at correctly surfacing the ObjectNotFound error with enough context. Recovery strategy is application-specific.
7.4.3 Distinguishing ObjectNotFound from a typo
A 404 on Get(<address>) for an address the application supplied directly (without going through a Manifest) is distinguishable from ObjectNotFound:
- The address may simply have never existed (typo, fabricated hash).
- This is a programming error, not a data-integrity event.
SDKs MAY surface this as a distinct error (AddressNotFound) or as ObjectNotFound with a "no Manifest context" marker. Implementation-defined.
7.3 Garbage and orphan collection
Orphan Objects accumulate over time. Sources include:
- Failed Append transactions (Phase 1 PUTs succeeded but Publish never happened).
- Aborted Publishes (Manifest PUT succeeded but Ref CAS lost; the staging writer either rebuilt and won, or abandoned, in either case orphaning the original Manifest).
- Reverted experiments (writer staged Phase 1 work, decided not to commit).
- Branches that were never merged or referenced again (
0008will detail).
At even modest write rates, orphan accumulation matters within months. DreamDB provides no protocol-level GC verb (it would require centralized coordination, conflicting with the lock-free design); operators run GC out-of-band. The recommended algorithm:
7.3.1 Two-step GC algorithm
Step 1: compute the reachable set.
Work scales with traversed metadata and reference entries, not only distinct
Object count: many Bucket records may share one VS Object. The reference
collector holds a live set and a per-physical-key expansion cache; this does
not guarantee that a billion-Object walk fits memory. An unreadable or
unauthenticated spatial dependency description aborts before the sweep, per
0007 §6.3.2. Missing entry-level sidecars never waive record-level VS edges.
This enumeration MUST be exhaustive over the Object kinds a Manifest can transitively reference. An implementation that encounters a content-hash reference it does not recognize MUST treat it as reachable (mark, do not sweep) rather than silently dropping it — the sweep deletes every non-reachable Object, so any gap in this list is silent data loss. (Implementations historically drifted out of sync here as new Object kinds were added in 0007/0010/0015/0020.)
Step 2: LIST + diff + age-threshold + DELETE.
7.3.2 Retention threshold and maintenance precondition
Operators SHOULD retain the default age threshold of 24 hours as an additional retention policy. Object age does not establish transaction liveness: a new transaction can reuse an old content-addressed Object, and a conformant repeated PUT may be a no-op or return 412 without refreshing Last-Modified (0005 §3.2).
Destructive GC MUST run only in an externally enforced maintenance window:
- Stop and drain all writers, staged transactions, uploads and Ref/root changes sharing the backend. Unpublished work that will resume MUST have its closure protected by an explicit retained Manifest root; otherwise finish or abort it before GC.
- Obtain a complete, settled Ref listing across the backend's Regions. Supply every Ref-less Manifest and historical snapshot still required by readers as an additional root.
- Keep the exclusion in force from before the mark phase until all deletions finish. Readers may continue only on retained snapshots.
- Resume publication after GC finishes. A timed-out or cancelled GC must have stopped all deletion activity before the fence is released.
The CLI requires --writers-paused for destructive runs; the Dataset GC API requires GcPolicy.writers_paused. These are explicit operator assertions, not lock acquisition or a protocol enforced on remote writers. The deployment must fence writes itself (for example by stopping all publishers and restricting backend write access). Without that capability, use dry-run only. Dry-run output collected concurrently is advisory, not a plan that may later be replayed as DELETEs.
An empty root set MUST refuse even in dry-run. --root <base32-manifest-hash> is repeatable and augments Ref roots; an unavailable explicit root or incomplete closure also refuses before sweep. Whole-backend destruction is a separate, explicitly authorized operation.
7.3.2.1 Why age and touch-to-extend do not permit concurrent GC
Counterexample: GC marks an old unreferenced Object O as an orphan; a short transaction re-PUTs O (412, unchanged timestamp) and publishes a new Manifest referencing O; GC then deletes O based on its old mark and age. Even transactions shorter than one hour admit this schedule.
Consequently neither bounded transaction duration, repeated PUT, a larger age threshold nor an extra Ref/HEAD check proves concurrent safety. The former touch-to-extend recommendation and unconditional concurrent-safety claim are withdrawn. An online reclamation protocol would need separately specified publication pins/epochs or equivalent coordination; it is not implemented by this maintenance-only collector.
7.3.3 GC operational notes
- GC MUST satisfy the maintenance preconditions above. It is not safe to run concurrently with publication; readers require explicitly retained roots.
- GC is idempotent — running it twice does no harm; the second run finds no additional candidates.
- GC SHOULD be run periodically (e.g., daily). Ad-hoc GC after large failed transactions is also valid.
- For multi-Region deployments, GC SHOULD walk Refs from every Region before deleting (a Ref in one Region might reference Objects another Region considers orphans).
This is the same "mark and sweep over an immutable substrate" pattern Git uses for unreachable commits.
7.3.4 Scaling: full-walk GC vs. incremental GC
The full-walk algorithm above is O(historical Manifests + Objects) in cost. A 10-year-old Space with live ingest at 1 Manifest/minute accumulates ~5M Manifests in history; the reachable-set walk traverses every one. For very long-running Spaces this becomes operationally expensive (hours to days per GC run, multi-GB working set).
Incremental GC (deferred to v0.1 as a fully-specified pattern) reduces this to amortized work proportional to recent activity:
- The operator periodically writes a GC checkpoint Object to the backend recording (a) the set of Manifest hashes reachable as of the checkpoint time, (b) the timestamp of the checkpoint.
- Subsequent GC runs walk only Manifests newer than the latest checkpoint plus a small fixed lookback window, take the union with the checkpoint's reachable set, then proceed with the LIST + diff phase as usual.
- Checkpoint Objects are themselves content-addressed and protected from deletion by being reachable from the GC tooling's own pseudo-Ref.
For v0, operators of Spaces with >100K historical Manifests SHOULD plan for either full-walk GC at low frequency (weekly / monthly) or a custom incremental approach. The full-walk algorithm is correct and conformant; it's just slow at extreme history depths. v0.1 will pin the incremental pattern as a normative SHOULD with checkpoint Object format.
8. Worked End-to-End Example
A 1-hour video is ingested, embedded for semantic search, and then queried.
8.1 Ingest
8.2 Query (cold-start, second day)
8.3 Stream (the moment of interest)
End-to-end query → playback < 200 ms. Sub-100 ms achievable with smaller modality buckets and PQ compression in 0007.
9. Out of Scope for this Document
- Cross-Timeline join verbs. DreamDB v0 has no such join (per
0001§11). Future spec MAY add one. - Push Subscribe / durable change streams. Watch (§4.1.2) remains pull-based; backend push and replayable event logs are out of scope.
- Branching verbs (
Branch,Merge). Manifest history walking and merge semantics are0008. - Garbage collection verb. GC is operator-level (per §7.3), not a protocol verb.
- Cross-Space federation. v0 SDK opens one Space per session. Multi-Space joins are application concerns.
10. Open Questions Surfaced by This Document
- OQ-33 (→ 0009 §7): Conformance test vectors for verb behavior. Resolved: full battery in
0009§7 covering standardAppend → Publishround-trip;Appendretry after mid-sequence failure; concurrentPublishreconciliation; cold-vs-hot-pathQuerylatency assertions; GC algorithm; Stream prefetch under simulated tail latency; Ref-freshness polling. - OQ-34: Resolved. §4.1.2 defines optional pull-based Ref Watch with explicit coalescing, recovery, cancellation and pinned-read semantics (#357). The decision is SDK composition, not a new backend notification verb; push Subscribe and durable event delivery are not promised.
- OQ-35 (→ 0008 §4.2): Explicit parent for
Publish. Resolved:Publish(session, manifest_spec={parents: [<hash>], ...})accepts an explicitparentsarray; absent → implicit parent = Session's loaded Manifest tip.
Next: 0007-streaming-encapsulation.md — defines the byte format inside Fragments, Spatial Buckets, and Time-bucketed batches; pins fragment durations, bucket-splitting thresholds, byte-range vs. inline storage decisions (OQ-23, OQ-24); and resolves OQ-4, OQ-7, OQ-13, OQ-14, OQ-15, OQ-20, OQ-21.