Spec 0012 — Federation & Cross-Cluster Queries
Status: Implemented (v0.2 read federation).
Depends on: spec/0001, spec/0002, spec/0005, spec/0008.
1. Purpose and boundary
Federation presents several immutable Dataset Manifests, stored on independent backends, as one read-only logical dataset. It does not add cross-backend transactions or a second data plane. A single coordinator stores the small Federation Manifest and its mutable Federation Ref; each logical shard remains an ordinary Dataset whose objects are read through an ordinary Connector.
v0.2 ships:
- a canonical Federation Manifest and address grammar;
- runtime binding from stable backend ids to Connectors;
- mirror failover within a logical shard;
- anchor-range pruning, time-range merge and vector top-K merge;
- an explicit distinction between availability quorum and complete coverage.
The following are not part of this format:
- backend URLs, credentials, capability tokens or retry deadlines;
- cross-backend writes, automatic rebalancing or distributed CAS;
- recursive Federation Manifests;
- out-of-core or automatically refreshed federation-level ANN indexes.
Those are deployment policy or later optimizations. They MUST NOT be smuggled
into the v1 maps as optional keys: every map below is CLOSED because it carries
reference and routing semantics (0002 §3.1.0).
2. Model
A logical shard is one partition of the logical data domain. It names one exact Dataset Manifest. A logical shard may have several replicas; every replica MUST serve the byte-identical closure of that Manifest. Replica count never increases data coverage and never counts more than once toward quorum.
The immutable manifest stores stable backend identifiers, not physical URLs.
At open time the caller supplies a resolver from each identifier to a
Connector. Changing an endpoint or rotating credentials therefore does not
change the Federation Manifest's content identity. Authentication and
authorization remain Connector responsibilities (0005 §8.4).
All shards carry one schema commitment: the BLAKE3 multihash of the canonical
CBOR representation stored under dreamdb.schema. A shard whose opened
Schema hashes differently is protocol corruption and MUST abort the whole
query; quorum never authorizes mixing incompatible score or payload domains.
3. Federation Manifest
Address:
3.1 CLOSED canonical CBOR
Every listed key occurs exactly once and no other key is allowed.
versionis exactly 1.parentsis strictly increasing by raw 33-byte multihash and has no duplicates.schemais the shared Schema hash defined in §2.shardsis non-empty and strictly increasing by UTF-8 bytes ofid.quorumis in1..=len(shards)and counts logical shards, not replicas.- shard and backend ids are 1–64 lowercase ASCII characters from
[a-z0-9_-], beginning with an alphanumeric character. replicasis non-empty and strictly increasing by UTF-8 bytes. Array order is canonical, not a latency preference; a runtime resolver may choose among available replicas but may not reinterpret them as distinct partitions.
Multihashes are 33-byte CBOR byte strings. Text base32 is diagnostic only.
3.1.1 Routed v2 snapshots
Version 2 has exactly the five v1 keys plus required router, a multihash of
a coordinator-local GraphIndex using the v5 inline routing profile (0013
§6). All v1 field rules remain unchanged. Version 1 MUST omit this key and
its bytes retain their meaning. Implementations lacking v2 MUST reject it
before querying or marking/sweeping; this is not an ignorable extension.
Before publication or query, verify router content hash and exact agreement
on shared Schema, every logical shard id and its immutable Manifest address.
Missing/stale/foreign/duplicate routes are malformed. The selected fanout MUST
be at least the manifest quorum: routing never lowers availability requirements.
Results name selected and intentionally pruned shards separately from failed
selected shards. partial=false does not certify global nearest-neighbor
recall. Ordinary v1 scatter-gather and anchor-range scope behavior are unchanged.
The router is durable before the Federation Ref moves and remains GC-reachable
through every retained Federation snapshot, including ancestors. Shard bindings
inside it name the same remote manifests as shards, not additional
coordinator-owned objects. A router decode/binding failure is an incomplete
closure and MUST prevent sweep. Opaque copies may preserve the bytes without
interpreting the router, under 0002 §3.1.0.
3.2 CLOSED scope
An unprunable shard uses:
An anchor-range shard uses:
The latter owns the half-open interval [start,end) and requires
start < end. Scope is a query-pruning assertion, not object authorization:
the child Manifest remains the only authority over the objects it names.
3.3 Parents and publication
A Federation Ref lives at:
Its body is exactly one 33-byte Federation Manifest hash. It is a mutable
RefStore path for timeout/retry classification just like refs/ (0005
§4 and §8.2). Publication writes and flushes the content-addressed manifest
before creating or CAS-advancing the ref. On advance, the old tip MUST be a
direct member of the new manifest's parents; a stale CAS is a publish
conflict. Re-publishing the exact current hash is idempotent.
There is one authoritative Federation Ref. v1 does not attempt multi-backend CAS.
4. Resolution
For each selected logical shard, a reader tries its resolved replica
Connectors until one can open the exact manifest hash. Each content-addressed
read is verified against that hash by the normal Dataset path.
The following are availability failures and permit trying another replica:
- no runtime binding for a backend id;
- Connector failure while opening or querying that replica.
A malformed object, Schema mismatch, unsupported format, or other protocol failure is not an availability vote. It aborts the whole query even if quorum could otherwise be met. Treating corruption as an absent shard would turn a safety boundary into an availability option.
5. Query semantics
Selected logical shards are evaluated concurrently. A time-range query skips
an anchor-range shard whose interval does not overlap the requested interval;
all shards are never pruned. For a pruned query the required success count is
min(quorum, selected_shards). Selecting no shard returns a complete empty
result.
5.1 Quorum is not completeness
If successful logical shards are fewer than the required count, the query fails. It MUST NOT return a value labelled partial.
If the required count is met, the query returns a value. It is complete only
when every selected logical shard succeeded. Otherwise it carries
partial=true and the ids of the failed logical shards. Thus a 9-of-10 result
with quorum=9 is available but incomplete; quorum never redefines the data
domain. A failed replica hidden by another working replica does not make the
logical-shard result partial.
This separation resolves OQ-48: v1 uses an absolute availability threshold; the returned completeness statement supplies the numerator/denominator fact without overloading the threshold.
5.2 Time-range merge
Each selected shard runs the ordinary Dataset time-range query. Rows are
merged in deterministic (time_anchor, shard_id) order and retain their shard
id. Equal anchors from different shards are distinct rows; federation does not
invent a cross-shard item identity.
5.3 Vector merge
Each selected shard returns its local min(K,n_s) best matches on the same
Schema-defined score scale. Concatenating those lists, sorting by
(score descending, shard_id, time_anchor), and retaining the first K
produces the global top-K of those shard results. A sparse shard with fewer
than K candidates returns all of them and is fully represented; it need not
fabricate K rows.
That statement is about deterministic merge, not ANN recall. Approximate search may omit true neighbors inside a shard. A fixed oversampling factor is a tuning choice and MUST NOT be presented as a universal recall proof. Scores from different Schema identities, algorithms, or compressor identities MUST NOT be mixed; §2's Schema commitment makes that state fail closed.
6. Reachability, GC and caches
A Federation Manifest is a coordinator-local root for its own parents; it is
not a magic remote GC root. Every child Dataset Manifest and its closure MUST
remain retained by a normal local Ref or another backend-specific retention
root for as long as a readable Federation Manifest names it. Publishing a
federation does not authorize deleting or moving a child's existing Ref.
This rule is deliberate: allowing a coordinator object to control a remote backend's collector would create cross-backend liveness and authorization that the Connector interface does not provide. An operator that copies a closure between backends uses ordinary verified GET/PUT operations and establishes a local retention root before advertising the replica.
Federation and Dataset Manifests are content addressed and may be cached by hash. Federation Refs and backend resolver bindings are mutable/runtime state and must not be cached as if content addressed.
7. Security boundary
Federation adds no credential format. Connectors authenticate to physical backends; their credentials, endpoints and expiry never enter canonical CBOR. A caller that can bind a backend id chooses where reads go and is therefore a trusted deployment component. Returned bytes are still verified against the manifest and object hashes, so a resolver cannot silently substitute different content under an existing federation identity.
Opaque whole-closure replication needs no federation-specific wire verb. A
semantic replicator must understand every reference-bearing format it walks;
an opaque replicator may copy an already enumerated path set and verify bytes
(0002 §3.1.0). OQ-40 is resolved on that basis: v1 does not standardize the
earlier draft's HTTP federate service endpoint.
7.1 Snapshot-bound operator copy plans (OQ-49)
The Rust copy_plan::CopyPlan API implements the versioned
dreamdb.copy-plan.v1 operator envelope. This is not a DreamDB Object
kind, registry extension, federation membership update, or remote GC authority.
Capability refusals below are explicit boundaries, not permission to emit
partial plans or claim that every valid Dataset format is supported.
The envelope is canonical CBOR with exactly kind (the profile string),
manifest (33-byte multihash), and objects (nonempty array). Each Object is
exactly [path: text, hash: bstr(33), size: uint]. Entries are strictly sorted
by backend-relative canonical path, with no duplicates; a path's terminal hash
MUST equal its declared hash. Ref paths are forbidden. The pinned Manifest
MUST appear. Limits are 100,000 Objects and 16 MiB of encoded plan. Endpoints,
credentials, source Ref names and destination root names are out-of-band.
A digest of the envelope proves byte identity, not closure completeness.
Export uses the existing Dataset semantic reachability walk extended with exact paths, not backend LIST plus hash matching. It traverses every named Dataset parent through genesis (maximum 10,000 Manifests), including retained lineage Tracks, and verifies every enumerated Object's complete bytes. A missing named parent is an error, not an implicit snapshot roll-up. Equal hashes at two required paths do not make either path dispensable.
The profile supports inline and paged ordinary Fragment (including packs/chunked Items), TimeBatch, Unbucketed, Constant, ScalarBucket, inline-vector and reference-vector SpatialBucket Tracks, their exact-vector sidecars, Genesis, entity key indexes, and registry SpatialIndex/ScalarIndex/VectorCompressor Objects. Scalar B-trees and their value-keyed bucket/pack Objects are also traversed. Typed and raw HotShards retain their exact cold binding, tombstone heads/pages retain their parent DAG, and backfill claims retain decided sets and basis Tracks across the copied Manifest history. TextIndex roots/posting pages, Vamana/Fresh graph pages and compressed/exact lineage, VideoItem time/key/fragment indexes and init segments, rendition playlists, and geometry descriptors/pages/payloads are also traversed at their exact physical keys.
Reference-mode VS records and Bucket-header SpatialIndex/VectorCompressor
dependencies use the same authenticated walk in copy planning and Dataset GC
(#365, 0007 §6.3.2). Operators MUST upgrade or fence older destination
collectors which omit those edges; successful byte copying alone cannot repair
an unsafe collector.
It explicitly refuses paged Manifest Track lists (distinct from supported paged Track indexes) and unsupported graph profiles. These remain valid Dataset formats; refusal denotes a missing consumer capability, not invalid data. This profile copies one Dataset snapshot; federation coordinator/child orchestration remains external. Unknown critical semantics MUST NOT be silently omitted.
Apply verifies existing destination bytes rather than trusting HEAD, key names, or a PUT's AlreadyExists result. It copies missing Objects with create-only PUTs and ensures the ordinary derived range-verification outboards exist. Those reconstructible outboards are not independent entries in this semantic plan. Before publishing a root it flushes, independently re-enumerates and verifies the destination snapshot, and requires exact equality with the plan. It then creates an ordinary destination Ref without retargeting another snapshot, verifies the resulting pointer, flushes, and reopens via that Ref. Only this successful path reports completion; copying a supplied list alone does not establish semantic completeness.
The operator MUST retain the source snapshot during transfer and exclude destination GC and mutation of the selected Ref through publication. The API's fence assertion acquires no distributed lock. Errors may leave reusable partial Objects; a late Ref/flush error may leave the root already published. Retrying the same plan/root is idempotent, while corrupt existing bytes or a different root target cause refusal, never overwrite. No failure path deletes data. The destination Ref retains its snapshot under the ordinary GC contract; preserving all copied history additionally requires the operator's ancestor retention policy or explicit roots. The plan file itself retains nothing. Whole Objects are buffered serially; this profile does not promise bounded memory for arbitrarily large Objects or a background replication service.
8. Deferred optimizations
- A federation-level GraphIndex may route vector queries to fewer logical
shards (
0013). It changes performance and recall, not the correctness of scatter/gather defined here. - Automatic shard rebalancing and read repair remain operator functions.
- Recursive Federation Manifests require a new wire version and explicit cycle and depth rules.
- Federated writes require an application-level coordinator; v1 handles reads only.
9. Conformance
dreamdb-conformance/vectors/0012/federation-manifest/ pins canonical bytes,
content hash and round-trip behavior. The Dataset integration check publishes,
reopens and queries two distinct filesystem roots through the public API; it
also distinguishes mirror failover, partial success and below-quorum refusal.
10. Open questions
- OQ-49: Resolved. §7.1 defines canonical snapshot-bound operator interchange, verified resumable copying and destination-root publication (#359). Explicit capability limits remain; reference-mode GC retention is implemented by #365.
- OQ-50: Whether coordinator tooling should audit that every child has a backend-local retention root.
- OQ-51: Resolved.
0013§6 defines exact-snapshot-bound router GraphIndexes, selected-shard traversal and refresh; the reference federation path implements that contract (#319). It does not promise global exact recall or automatic representative extraction. - OQ-52: Runtime mirror discovery. It cannot change Federation Manifest identity or weaken explicit resolver authorization.