DreamDB Specification — 0022: Fragment Packs
Status: Normative (v0.1), 2026-05-22. Implemented in the reference impl (
dreamdb-protocoltrackFragmentEntry6-tuplepack_offset;dreamdb-datasetappendpacking path) and gated by conformance: the packed-entry byte format at0009§8.8 (dreamdb-conformance/vectors/0022/: multi-item pack + two-packs-in-one-track roundtrip, plus the 6-tuple vector atvectors/0007/track-object/003), and the behavioral pack/byte-range/mixed-field assertions indreamdb-dataset/tests/fragment_pack.rs. Builds on0001,0002§7.3,0007§5–§6,0014(item chunking).
1. Purpose
0007 §5 defines the Fragment Object kind for media items (image, audio, video). The historical convention is one Fragment Object per Item — a single PUT per item on the write path. This is correct, but for small items (e.g. 16–64 KB JPEGs typical of a CLIP-encoded image corpus) and a write path that has to traverse the WAN to reach object storage, the per-PUT overhead (TCP/TLS handshake + RTT + HTTP framing) dominates over the actual payload bandwidth.
A Fragment Pack is a single Fragment Object that bundles many items together. The Track Object's object_index carries one FragmentEntry per item with a pack_offset field locating the item's bytes within the pack. Readers do byte-range GETs to extract individual items from a pack.
The net effect at write time: N items become ⌈N/pack_items⌉ S3 PUTs instead of N. At read time: nothing changes for sequential scans (they fetch the pack once anyway); random-access reads pay one ranged GET per item, the same RTT as today's one-object-per-item scheme.
This document fixes the on-wire format of FragmentEntry's pack_offset field, the pack body format (such as it is — see §3), and the read/write obligations.
2. Why no pack header
The simplest possible design: the pack body is just the concatenation of item bytes. Per-item indexing lives entirely in the Track Object's FragmentEntry (extended with pack_offset), not in the pack itself.
Rationale:
- Track Object is already the source of truth for per-item placement (
0005§5.3.1, Manifest Supremacy). Putting a duplicate index in the pack header would mean two sources of truth. - Sequential reads stream the pack and iterate its known entries; a header would be pure overhead.
- Random-access reads (rare for image corpora) pay one ranged GET, same as today's one-object-per-item scheme.
- No pack-header version field means no pack-header format-evolution problem.
The downside is that a pack is only meaningful in conjunction with its FragmentEntries. A pack Object cannot self-describe its items. This is a deliberate trade: pack Objects without their owning Track are unrecoverable data, but DreamDB's content-addressing already requires the Track to resolve any item, so the constraint adds nothing.
3. On-wire format
3.1 Pack body
A FragmentPack Object's body is raw byte concatenation of item payloads in t_start order:
No prefix. No header. No alignment. Each item begins immediately after the previous item ends.
3.2 Address
Packs are stored under the same address namespace as ordinary Fragment Objects (0007 §5.1): <timeline>/<modality>/<time-bucket>/<hash>. The content hash uniquely identifies the pack; no separate pack path prefix.
A reader cannot distinguish a pack from a single-item Fragment by looking at the Object alone — the distinction lives in the FragmentEntry that references it.
3.3 FragmentEntry extension
Per 0007 §7.3.1, FragmentEntry is a positional CBOR array:
- 4 elements:
[t_start, t_end, byte_size, fragment_address]— historical single-Fragment entry. - 5 elements:
[t_start, t_end, byte_size, fragment_address, true]— chunked manifest (per0014). - 6 elements:
[t_start, t_end, byte_size, fragment_address, false, pack_offset]— packed entry (this spec).
Where:
byte_sizeis the size of THIS item within the pack (not the size of the whole pack).fragment_addressis the content hash of the pack Object (multiple entries may share the same value).- The 5th element is the existing
is_manifestflag, which MUST befalsefor packed entries (chunking and packing are mutually exclusive). pack_offsetis the byte offset of this item's bytes within the pack body. Encoded as a CBOR unsigned integer (≤ u32::MAX).
A reader retrieves the item's bytes by issuing GET fragment_address [pack_offset .. pack_offset + byte_size].
Compatibility classification. pack_offset is a critical extension under 0002 §3.1.0, on condition 1 (payload interpretation) alone. An implementation that ignores the 6th element performs a whole-Object GET and returns the entire pack body in place of the single Item, silently, for every entry in the pack.
It is not critical on condition 3: a packed entry introduces no new downstream reference. The pack Object's address is fragment_address, which every collector already marks, and nothing inside the pack body needs traversal. Packing adds no marking obligation.
These two statements are independent and must not be collapsed. "No new reference" is a statement about the reference closure; it is not a statement that collectors are unaffected by the format. Should packed Tracks ever be placed behind a new discriminant, reader, writer and collector must each either support that discriminant or refuse explicitly — a collector that does not recognise it will refuse at type resolution, which is the correct outcome but is still a hard runtime dependency that an upgrade plan must account for.
pack_offset is a named historical exception under 0002 §3.1.0.1; per-role obligations and the rollback rule are in 0002 §3.1.0.3.
3.4 Pack invariants
For all FragmentEntries that share a fragment_address (i.e., reference the same pack):
- The entries'
[pack_offset, pack_offset + byte_size)ranges MUST be disjoint and contiguous, starting at offset 0. - The pack's body size MUST equal the sum of the entries'
byte_sizevalues (no gaps, no trailing bytes). - Item time ranges MAY overlap (Items at the same anchor across modalities is allowed); the pack stores items in the time order they were emitted.
Writers MUST observe (1) and (2). Readers MAY assume (1) and (2) and MUST emit a clear error if a pack's actual byte length does not match the sum of entries.
4. Writer obligations
When Schema.<field>.pack_items = Some(N) and N > 1 and chunk_size = None:
- Collect items for this modality from one or more samples in the current
append_manycall. - Group items into chunks of ≤ N items each. Within each chunk:
- Compute the pack body = concatenation of item bytes in collection order.
- PUT the pack as a single Object at
<timeline>/<modality>/<time-bucket=0>/<content-hash>. - For each item in the chunk, emit a
FragmentEntrywithpack_offsetset to the item's byte offset andbyte_sizeset to the item's length.
- Append all emitted
FragmentEntryinstances to the Track Object'sobject_index(alongside any pre-existing entries; per0007§6.6, multi-entry-per-cell is the LSM steady state).
Writers MAY parallelize pack PUTs across chunks. The recommended concurrency is 64-wide (the same as for individual Fragment PUTs), implemented via buffer_unordered or equivalent.
When pack_items = None, pack_items = Some(1), or chunk_size = Some(_):
- Writers MUST emit single-item Fragments (the historical path). The 4-element or 5-element FragmentEntry shapes apply.
5. Reader obligations
When reading a FragmentEntry:
- If
is_manifest = true: follow the ItemManifest path (0014). - Else if the entry has 6 positional elements with the 6th being a CBOR unsigned integer: this is a packed entry. Fetch the item by
GET fragment_address [pack_offset .. pack_offset + byte_size]. - Else: this is a single-Fragment entry. Fetch the item by
GET fragment_address(whole-Object).
A reader that does not implement step 2 MUST NOT fall through to step 3 for a packed entry: that path returns the whole pack body as the Item. Per 0002 §3.1.0 it MUST refuse explicitly instead.
Readers MAY cache pack bodies after a full GET to amortize subsequent ranged reads from the same pack within the same scan. Implementations SHOULD do this when the read iterator is sequential and item entries are consecutive (a common pattern for training-data streaming).
6. Compaction interaction
Packs are immutable like all DreamDB Objects. The bucket compactor in 0021
targets the capacity floor K, not F=1, and is not a Fragment repacker.
Dataset::repack_fragments(field, limits) is an explicit, whole-field operation
for ordinary Image/Audio/Video Fragment Tracks. It does not support chunked Items,
CMAF Tracks or VideoItem Tracks; those refuse before writes. It uses the existing
6-tuple encoding, preserves every entry's anchor interval and Item bytes, and
changes only the selected field's active Track in one ordinary Manifest/Ref CAS.
Physical Fragment addresses/offsets may change; callers must resolve locators
against the selected snapshot. Old snapshots and their content remain readable.
The operator supplies positive max_items, max_input_bytes and
max_pack_bytes limits. Planning enumerates the field's Track metadata first;
the item limit, summed Item bytes and summed physical input Object sizes are checked before reading
payloads or writing. No unbounded-payload or constant-metadata-memory claim is
made. A field over the input limits refuses; v1 does not silently repack a prefix.
Output groups follow Track order and contain at most Schema pack_items Items
and max_pack_bytes bytes (the latter is at most u32::MAX). An individual Item
larger than that target remains a single unpacked Fragment; it is neither split
nor rejected for exceeding the output target, but still counts against input
limits. Packing must be enabled in Schema; this operation does not change Schema.
All reads, slicing and child-Manifest validation precede writes. Repacking publishes only if the candidate has strictly fewer distinct current-root Fragment addresses. Otherwise it is a zero-write no-op, including repeated execution on its own output. Different legal packing geometry alone does not justify rewriting. Normal stale-parent/CAS rules apply; no automatic retry retargets a newer snapshot. No GC is invoked: retained parents, other Refs and pinned snapshots can keep all old Objects alive. Reduced current-root count is not reclaimed storage.
7. Why this design (vs alternatives)
| Alternative | Why rejected |
|---|---|
| Pack header CBOR with per-item index inside the pack | Duplicates the Track Object's index; two sources of truth; format-evolution problem. |
Separate pack/ path prefix at the address layer | Adds visible path complexity; consumer needs to know two paths exist. Content-addressing already disambiguates. |
New InlineObjectIndex::FragmentPack variant in Track Object | Adds a new variant to 129+ match sites across the codebase. Not justified for what is a small per-entry shape change. |
| Pack with internal byte-range index AND Track-level pack_offset | Strictly more bytes for the same information. |
Chunking (0014) for packing | Chunking splits ONE big item into N PUTs. Packing bundles N small items into ONE PUT. Opposite directions; can't compose. |
8. Conformance requirements (catalogued in 0009)
The conformance categories for fragment packs are catalogued at 0009 §8.8:
- Round-trip a pack: write N=8 items with pack_items=4 → expect 2 packs in S3 → read back, byte-equality vs originals.
- Mixed packs in one Track: append in 3 batches with different N each batch → expect 3 separate packs, item count matches.
- Mixed packed and unpacked Fields in one Schema: confirm per-Field independence of the
pack_itemsknob. - Reject pack + chunk combination: ingest a sample with both
chunk_sizeandpack_items > 1set on the same field → writer MUST refuse, OR writer MUST silently degrade to chunking (current behavior). - Byte-range fetch: a single-item GET with
pack_offsetset returns exactly[pack_offset .. pack_offset + byte_size]from the pack.
9. Measured impact
10K imagenet S3 ingest, laptop → us-east-1, fair A/B (back-to-back warm runs, 2026-05-22):
| Mode | S3 image objects | Wall-clock | Rate |
|---|---|---|---|
| Unpacked (1 PUT/item) | 10,000 | 72.0 s | 138.9 /s |
| Packed (pack_items=32) | 313 | 66.1 s | 151.3 /s |
The packing reduces S3 PUT count 32×. Wall-clock improves ~8% — modest because CLIP-MPS at ~150 /s is the dominant cost on this hardware. On GPU hardware where CLIP encoding is no longer the floor, the network savings translate to a much larger wall-clock win (estimated 5–10× per design/0008 §"Expected payoff").
The structural change is required for billion-scale ingest where individual-item PUTs run into S3 request-rate limits and dollar cost. With packing, network cost scales as O(N / pack_items) rather than O(N).
10. Open questions
The #345 public-boundary workload (64 distinct 1 KiB Items, pack_items=8, one 64-item scan batch, in-memory backend) measured:
| Append batch | Current-root packs | Payload bytes | Pack bytes min/max | Scan pack GETs |
|---|---|---|---|---|
| 1 Item | 64 | 65536 | 1024 / 1024 | 64 |
| 64 Items | 8 | 65536 | 8192 / 8192 | 8 |
Counts exclude metadata and retained historical roots, and distinct payloads exclude content deduplication as the explanation. Explicit repacking with an 8192-byte output target reduces the first case to 8 current-root packs, preserving both new and historical reads. This is request-count evidence, not a WAN timing, cloud-cost, universal byte optimum or physical-storage-reclamation measurement.
| OQ | Description |
|---|---|
| OQ-97 | RESOLVED — schema-pinned. pack_items is a persisted Schema field, declared per field at schema-construction time and serialised into the Schema Object (dreamdb-dataset schema.rs; exposed that way by the Python bindings' field constructors). append_many takes no packing parameter and no per-call override exists, so the packing geometry of a field is a property of the dataset rather than of whoever happens to be writing to it. |
| OQ-98 | RESOLVED — workload/operator policy, not a universal optimum. Schema pack_items remains the count ceiling; explicit repacking also accepts a bounded caller byte target (§6). The measurements above justify neither a default 2 MB target nor a universal performance optimum. |
| OQ-99 | RESOLVED — explicit bounded repacking, §6, after observed current-root pack/request sprawl. No automatic execution, GC or historical rewrite. |
Future changes to default sizing require their own workload evidence.