DreamDB

Spec 0012 — Federation & Cross-Cluster Queries

Status: Implemented (v0.2 read federation). Depends on: spec/0001, spec/0002, spec/0005, spec/0008.


1. Purpose and boundary

Federation presents several immutable Dataset Manifests, stored on independent backends, as one read-only logical dataset. It does not add cross-backend transactions or a second data plane. A single coordinator stores the small Federation Manifest and its mutable Federation Ref; each logical shard remains an ordinary Dataset whose objects are read through an ordinary Connector.

v0.2 ships:

  • a canonical Federation Manifest and address grammar;
  • runtime binding from stable backend ids to Connectors;
  • mirror failover within a logical shard;
  • anchor-range pruning, time-range merge and vector top-K merge;
  • an explicit distinction between availability quorum and complete coverage.

The following are not part of this format:

  • backend URLs, credentials, capability tokens or retry deadlines;
  • cross-backend writes, automatic rebalancing or distributed CAS;
  • recursive Federation Manifests;
  • out-of-core or automatically refreshed federation-level ANN indexes.

Those are deployment policy or later optimizations. They MUST NOT be smuggled into the v1 maps as optional keys: every map below is CLOSED because it carries reference and routing semantics (0002 §3.1.0).

2. Model

A logical shard is one partition of the logical data domain. It names one exact Dataset Manifest. A logical shard may have several replicas; every replica MUST serve the byte-identical closure of that Manifest. Replica count never increases data coverage and never counts more than once toward quorum.

The immutable manifest stores stable backend identifiers, not physical URLs. At open time the caller supplies a resolver from each identifier to a Connector. Changing an endpoint or rotating credentials therefore does not change the Federation Manifest's content identity. Authentication and authorization remain Connector responsibilities (0005 §8.4).

All shards carry one schema commitment: the BLAKE3 multihash of the canonical CBOR representation stored under dreamdb.schema. A shard whose opened Schema hashes differently is protocol corruption and MUST abort the whole query; quorum never authorizes mixing incompatible score or payload domains.

3. Federation Manifest

Address:

federation-manifests/<multihash-of-canonical-CBOR-bytes>

3.1 CLOSED canonical CBOR

{
  "version": 1,
  "parents": [<multihash>, ...],
  "schema": <multihash>,
  "shards": [
    {
      "id": <text>,
      "manifest": <multihash>,
      "replicas": [<backend-id>, ...],
      "scope": <scope>
    }, ...
  ],
  "quorum": <u32>
}

Every listed key occurs exactly once and no other key is allowed.

  • version is exactly 1.
  • parents is strictly increasing by raw 33-byte multihash and has no duplicates.
  • schema is the shared Schema hash defined in §2.
  • shards is non-empty and strictly increasing by UTF-8 bytes of id.
  • quorum is in 1..=len(shards) and counts logical shards, not replicas.
  • shard and backend ids are 1–64 lowercase ASCII characters from [a-z0-9_-], beginning with an alphanumeric character.
  • replicas is non-empty and strictly increasing by UTF-8 bytes. Array order is canonical, not a latency preference; a runtime resolver may choose among available replicas but may not reinterpret them as distinct partitions.

Multihashes are 33-byte CBOR byte strings. Text base32 is diagnostic only.

3.1.1 Routed v2 snapshots

Version 2 has exactly the five v1 keys plus required router, a multihash of a coordinator-local GraphIndex using the v5 inline routing profile (0013 §6). All v1 field rules remain unchanged. Version 1 MUST omit this key and its bytes retain their meaning. Implementations lacking v2 MUST reject it before querying or marking/sweeping; this is not an ignorable extension.

Before publication or query, verify router content hash and exact agreement on shared Schema, every logical shard id and its immutable Manifest address. Missing/stale/foreign/duplicate routes are malformed. The selected fanout MUST be at least the manifest quorum: routing never lowers availability requirements. Results name selected and intentionally pruned shards separately from failed selected shards. partial=false does not certify global nearest-neighbor recall. Ordinary v1 scatter-gather and anchor-range scope behavior are unchanged.

The router is durable before the Federation Ref moves and remains GC-reachable through every retained Federation snapshot, including ancestors. Shard bindings inside it name the same remote manifests as shards, not additional coordinator-owned objects. A router decode/binding failure is an incomplete closure and MUST prevent sweep. Opaque copies may preserve the bytes without interpreting the router, under 0002 §3.1.0.

3.2 CLOSED scope

An unprunable shard uses:

{ "kind": "all" }

An anchor-range shard uses:

{ "kind": "anchor-range", "start": <u64>, "end": <u64> }

The latter owns the half-open interval [start,end) and requires start < end. Scope is a query-pruning assertion, not object authorization: the child Manifest remains the only authority over the objects it names.

3.3 Parents and publication

A Federation Ref lives at:

federation-refs/<ref-name>

Its body is exactly one 33-byte Federation Manifest hash. It is a mutable RefStore path for timeout/retry classification just like refs/ (0005 §4 and §8.2). Publication writes and flushes the content-addressed manifest before creating or CAS-advancing the ref. On advance, the old tip MUST be a direct member of the new manifest's parents; a stale CAS is a publish conflict. Re-publishing the exact current hash is idempotent.

There is one authoritative Federation Ref. v1 does not attempt multi-backend CAS.

4. Resolution

For each selected logical shard, a reader tries its resolved replica Connectors until one can open the exact manifest hash. Each content-addressed read is verified against that hash by the normal Dataset path.

The following are availability failures and permit trying another replica:

  • no runtime binding for a backend id;
  • Connector failure while opening or querying that replica.

A malformed object, Schema mismatch, unsupported format, or other protocol failure is not an availability vote. It aborts the whole query even if quorum could otherwise be met. Treating corruption as an absent shard would turn a safety boundary into an availability option.

5. Query semantics

Selected logical shards are evaluated concurrently. A time-range query skips an anchor-range shard whose interval does not overlap the requested interval; all shards are never pruned. For a pruned query the required success count is min(quorum, selected_shards). Selecting no shard returns a complete empty result.

5.1 Quorum is not completeness

If successful logical shards are fewer than the required count, the query fails. It MUST NOT return a value labelled partial.

If the required count is met, the query returns a value. It is complete only when every selected logical shard succeeded. Otherwise it carries partial=true and the ids of the failed logical shards. Thus a 9-of-10 result with quorum=9 is available but incomplete; quorum never redefines the data domain. A failed replica hidden by another working replica does not make the logical-shard result partial.

This separation resolves OQ-48: v1 uses an absolute availability threshold; the returned completeness statement supplies the numerator/denominator fact without overloading the threshold.

5.2 Time-range merge

Each selected shard runs the ordinary Dataset time-range query. Rows are merged in deterministic (time_anchor, shard_id) order and retain their shard id. Equal anchors from different shards are distinct rows; federation does not invent a cross-shard item identity.

5.3 Vector merge

Each selected shard returns its local min(K,n_s) best matches on the same Schema-defined score scale. Concatenating those lists, sorting by (score descending, shard_id, time_anchor), and retaining the first K produces the global top-K of those shard results. A sparse shard with fewer than K candidates returns all of them and is fully represented; it need not fabricate K rows.

That statement is about deterministic merge, not ANN recall. Approximate search may omit true neighbors inside a shard. A fixed oversampling factor is a tuning choice and MUST NOT be presented as a universal recall proof. Scores from different Schema identities, algorithms, or compressor identities MUST NOT be mixed; §2's Schema commitment makes that state fail closed.

6. Reachability, GC and caches

A Federation Manifest is a coordinator-local root for its own parents; it is not a magic remote GC root. Every child Dataset Manifest and its closure MUST remain retained by a normal local Ref or another backend-specific retention root for as long as a readable Federation Manifest names it. Publishing a federation does not authorize deleting or moving a child's existing Ref.

This rule is deliberate: allowing a coordinator object to control a remote backend's collector would create cross-backend liveness and authorization that the Connector interface does not provide. An operator that copies a closure between backends uses ordinary verified GET/PUT operations and establishes a local retention root before advertising the replica.

Federation and Dataset Manifests are content addressed and may be cached by hash. Federation Refs and backend resolver bindings are mutable/runtime state and must not be cached as if content addressed.

7. Security boundary

Federation adds no credential format. Connectors authenticate to physical backends; their credentials, endpoints and expiry never enter canonical CBOR. A caller that can bind a backend id chooses where reads go and is therefore a trusted deployment component. Returned bytes are still verified against the manifest and object hashes, so a resolver cannot silently substitute different content under an existing federation identity.

Opaque whole-closure replication needs no federation-specific wire verb. A semantic replicator must understand every reference-bearing format it walks; an opaque replicator may copy an already enumerated path set and verify bytes (0002 §3.1.0). OQ-40 is resolved on that basis: v1 does not standardize the earlier draft's HTTP federate service endpoint.

7.1 Snapshot-bound operator copy plans (OQ-49)

The Rust copy_plan::CopyPlan API implements the versioned dreamdb.copy-plan.v1 operator envelope. This is not a DreamDB Object kind, registry extension, federation membership update, or remote GC authority. Capability refusals below are explicit boundaries, not permission to emit partial plans or claim that every valid Dataset format is supported.

The envelope is canonical CBOR with exactly kind (the profile string), manifest (33-byte multihash), and objects (nonempty array). Each Object is exactly [path: text, hash: bstr(33), size: uint]. Entries are strictly sorted by backend-relative canonical path, with no duplicates; a path's terminal hash MUST equal its declared hash. Ref paths are forbidden. The pinned Manifest MUST appear. Limits are 100,000 Objects and 16 MiB of encoded plan. Endpoints, credentials, source Ref names and destination root names are out-of-band. A digest of the envelope proves byte identity, not closure completeness.

Export uses the existing Dataset semantic reachability walk extended with exact paths, not backend LIST plus hash matching. It traverses every named Dataset parent through genesis (maximum 10,000 Manifests), including retained lineage Tracks, and verifies every enumerated Object's complete bytes. A missing named parent is an error, not an implicit snapshot roll-up. Equal hashes at two required paths do not make either path dispensable.

The profile supports inline and paged ordinary Fragment (including packs/chunked Items), TimeBatch, Unbucketed, Constant, ScalarBucket, inline-vector and reference-vector SpatialBucket Tracks, their exact-vector sidecars, Genesis, entity key indexes, and registry SpatialIndex/ScalarIndex/VectorCompressor Objects. Scalar B-trees and their value-keyed bucket/pack Objects are also traversed. Typed and raw HotShards retain their exact cold binding, tombstone heads/pages retain their parent DAG, and backfill claims retain decided sets and basis Tracks across the copied Manifest history. TextIndex roots/posting pages, Vamana/Fresh graph pages and compressed/exact lineage, VideoItem time/key/fragment indexes and init segments, rendition playlists, and geometry descriptors/pages/payloads are also traversed at their exact physical keys.

Reference-mode VS records and Bucket-header SpatialIndex/VectorCompressor dependencies use the same authenticated walk in copy planning and Dataset GC (#365, 0007 §6.3.2). Operators MUST upgrade or fence older destination collectors which omit those edges; successful byte copying alone cannot repair an unsafe collector.

It explicitly refuses paged Manifest Track lists (distinct from supported paged Track indexes) and unsupported graph profiles. These remain valid Dataset formats; refusal denotes a missing consumer capability, not invalid data. This profile copies one Dataset snapshot; federation coordinator/child orchestration remains external. Unknown critical semantics MUST NOT be silently omitted.

Apply verifies existing destination bytes rather than trusting HEAD, key names, or a PUT's AlreadyExists result. It copies missing Objects with create-only PUTs and ensures the ordinary derived range-verification outboards exist. Those reconstructible outboards are not independent entries in this semantic plan. Before publishing a root it flushes, independently re-enumerates and verifies the destination snapshot, and requires exact equality with the plan. It then creates an ordinary destination Ref without retargeting another snapshot, verifies the resulting pointer, flushes, and reopens via that Ref. Only this successful path reports completion; copying a supplied list alone does not establish semantic completeness.

The operator MUST retain the source snapshot during transfer and exclude destination GC and mutation of the selected Ref through publication. The API's fence assertion acquires no distributed lock. Errors may leave reusable partial Objects; a late Ref/flush error may leave the root already published. Retrying the same plan/root is idempotent, while corrupt existing bytes or a different root target cause refusal, never overwrite. No failure path deletes data. The destination Ref retains its snapshot under the ordinary GC contract; preserving all copied history additionally requires the operator's ancestor retention policy or explicit roots. The plan file itself retains nothing. Whole Objects are buffered serially; this profile does not promise bounded memory for arbitrarily large Objects or a background replication service.

8. Deferred optimizations

  • A federation-level GraphIndex may route vector queries to fewer logical shards (0013). It changes performance and recall, not the correctness of scatter/gather defined here.
  • Automatic shard rebalancing and read repair remain operator functions.
  • Recursive Federation Manifests require a new wire version and explicit cycle and depth rules.
  • Federated writes require an application-level coordinator; v1 handles reads only.

9. Conformance

dreamdb-conformance/vectors/0012/federation-manifest/ pins canonical bytes, content hash and round-trip behavior. The Dataset integration check publishes, reopens and queries two distinct filesystem roots through the public API; it also distinguishes mirror failover, partial success and below-quorum refusal.

10. Open questions

  • OQ-49: Resolved. §7.1 defines canonical snapshot-bound operator interchange, verified resumable copying and destination-root publication (#359). Explicit capability limits remain; reference-mode GC retention is implemented by #365.
  • OQ-50: Whether coordinator tooling should audit that every child has a backend-local retention root.
  • OQ-51: Resolved. 0013 §6 defines exact-snapshot-bound router GraphIndexes, selected-shard traversal and refresh; the reference federation path implements that contract (#319). It does not promise global exact recall or automatic representative extraction.
  • OQ-52: Runtime mirror discovery. It cannot change Federation Manifest identity or weaken explicit resolver authorization.