DreamDB

DreamDB Specification — 0022: Fragment Packs

Status: Normative (v0.1), 2026-05-22. Implemented in the reference impl (dreamdb-protocol track FragmentEntry 6-tuple pack_offset; dreamdb-dataset append packing path) and gated by conformance: the packed-entry byte format at 0009 §8.8 (dreamdb-conformance/vectors/0022/: multi-item pack + two-packs-in-one-track roundtrip, plus the 6-tuple vector at vectors/0007/track-object/003), and the behavioral pack/byte-range/mixed-field assertions in dreamdb-dataset/tests/fragment_pack.rs. Builds on 0001, 0002 §7.3, 0007 §5–§6, 0014 (item chunking).

1. Purpose

0007 §5 defines the Fragment Object kind for media items (image, audio, video). The historical convention is one Fragment Object per Item — a single PUT per item on the write path. This is correct, but for small items (e.g. 16–64 KB JPEGs typical of a CLIP-encoded image corpus) and a write path that has to traverse the WAN to reach object storage, the per-PUT overhead (TCP/TLS handshake + RTT + HTTP framing) dominates over the actual payload bandwidth.

A Fragment Pack is a single Fragment Object that bundles many items together. The Track Object's object_index carries one FragmentEntry per item with a pack_offset field locating the item's bytes within the pack. Readers do byte-range GETs to extract individual items from a pack.

The net effect at write time: N items become ⌈N/pack_items⌉ S3 PUTs instead of N. At read time: nothing changes for sequential scans (they fetch the pack once anyway); random-access reads pay one ranged GET per item, the same RTT as today's one-object-per-item scheme.

This document fixes the on-wire format of FragmentEntry's pack_offset field, the pack body format (such as it is — see §3), and the read/write obligations.

2. Why no pack header

The simplest possible design: the pack body is just the concatenation of item bytes. Per-item indexing lives entirely in the Track Object's FragmentEntry (extended with pack_offset), not in the pack itself.

Rationale:

  • Track Object is already the source of truth for per-item placement (0005 §5.3.1, Manifest Supremacy). Putting a duplicate index in the pack header would mean two sources of truth.
  • Sequential reads stream the pack and iterate its known entries; a header would be pure overhead.
  • Random-access reads (rare for image corpora) pay one ranged GET, same as today's one-object-per-item scheme.
  • No pack-header version field means no pack-header format-evolution problem.

The downside is that a pack is only meaningful in conjunction with its FragmentEntries. A pack Object cannot self-describe its items. This is a deliberate trade: pack Objects without their owning Track are unrecoverable data, but DreamDB's content-addressing already requires the Track to resolve any item, so the constraint adds nothing.

3. On-wire format

3.1 Pack body

A FragmentPack Object's body is raw byte concatenation of item payloads in t_start order:

[item_0 bytes][item_1 bytes][item_2 bytes]...[item_{N-1} bytes]

No prefix. No header. No alignment. Each item begins immediately after the previous item ends.

3.2 Address

Packs are stored under the same address namespace as ordinary Fragment Objects (0007 §5.1): <timeline>/<modality>/<time-bucket>/<hash>. The content hash uniquely identifies the pack; no separate pack path prefix.

A reader cannot distinguish a pack from a single-item Fragment by looking at the Object alone — the distinction lives in the FragmentEntry that references it.

3.3 FragmentEntry extension

Per 0007 §7.3.1, FragmentEntry is a positional CBOR array:

  • 4 elements: [t_start, t_end, byte_size, fragment_address] — historical single-Fragment entry.
  • 5 elements: [t_start, t_end, byte_size, fragment_address, true] — chunked manifest (per 0014).
  • 6 elements: [t_start, t_end, byte_size, fragment_address, false, pack_offset] — packed entry (this spec).

Where:

  • byte_size is the size of THIS item within the pack (not the size of the whole pack).
  • fragment_address is the content hash of the pack Object (multiple entries may share the same value).
  • The 5th element is the existing is_manifest flag, which MUST be false for packed entries (chunking and packing are mutually exclusive).
  • pack_offset is the byte offset of this item's bytes within the pack body. Encoded as a CBOR unsigned integer (≤ u32::MAX).

A reader retrieves the item's bytes by issuing GET fragment_address [pack_offset .. pack_offset + byte_size].

Compatibility classification. pack_offset is a critical extension under 0002 §3.1.0, on condition 1 (payload interpretation) alone. An implementation that ignores the 6th element performs a whole-Object GET and returns the entire pack body in place of the single Item, silently, for every entry in the pack.

It is not critical on condition 3: a packed entry introduces no new downstream reference. The pack Object's address is fragment_address, which every collector already marks, and nothing inside the pack body needs traversal. Packing adds no marking obligation.

These two statements are independent and must not be collapsed. "No new reference" is a statement about the reference closure; it is not a statement that collectors are unaffected by the format. Should packed Tracks ever be placed behind a new discriminant, reader, writer and collector must each either support that discriminant or refuse explicitly — a collector that does not recognise it will refuse at type resolution, which is the correct outcome but is still a hard runtime dependency that an upgrade plan must account for.

pack_offset is a named historical exception under 0002 §3.1.0.1; per-role obligations and the rollback rule are in 0002 §3.1.0.3.

3.4 Pack invariants

For all FragmentEntries that share a fragment_address (i.e., reference the same pack):

  1. The entries' [pack_offset, pack_offset + byte_size) ranges MUST be disjoint and contiguous, starting at offset 0.
  2. The pack's body size MUST equal the sum of the entries' byte_size values (no gaps, no trailing bytes).
  3. Item time ranges MAY overlap (Items at the same anchor across modalities is allowed); the pack stores items in the time order they were emitted.

Writers MUST observe (1) and (2). Readers MAY assume (1) and (2) and MUST emit a clear error if a pack's actual byte length does not match the sum of entries.

4. Writer obligations

When Schema.<field>.pack_items = Some(N) and N > 1 and chunk_size = None:

  1. Collect items for this modality from one or more samples in the current append_many call.
  2. Group items into chunks of ≤ N items each. Within each chunk:
    • Compute the pack body = concatenation of item bytes in collection order.
    • PUT the pack as a single Object at <timeline>/<modality>/<time-bucket=0>/<content-hash>.
    • For each item in the chunk, emit a FragmentEntry with pack_offset set to the item's byte offset and byte_size set to the item's length.
  3. Append all emitted FragmentEntry instances to the Track Object's object_index (alongside any pre-existing entries; per 0007 §6.6, multi-entry-per-cell is the LSM steady state).

Writers MAY parallelize pack PUTs across chunks. The recommended concurrency is 64-wide (the same as for individual Fragment PUTs), implemented via buffer_unordered or equivalent.

When pack_items = None, pack_items = Some(1), or chunk_size = Some(_):

  • Writers MUST emit single-item Fragments (the historical path). The 4-element or 5-element FragmentEntry shapes apply.

5. Reader obligations

When reading a FragmentEntry:

  1. If is_manifest = true: follow the ItemManifest path (0014).
  2. Else if the entry has 6 positional elements with the 6th being a CBOR unsigned integer: this is a packed entry. Fetch the item by GET fragment_address [pack_offset .. pack_offset + byte_size].
  3. Else: this is a single-Fragment entry. Fetch the item by GET fragment_address (whole-Object).

A reader that does not implement step 2 MUST NOT fall through to step 3 for a packed entry: that path returns the whole pack body as the Item. Per 0002 §3.1.0 it MUST refuse explicitly instead.

Readers MAY cache pack bodies after a full GET to amortize subsequent ranged reads from the same pack within the same scan. Implementations SHOULD do this when the read iterator is sequential and item entries are consecutive (a common pattern for training-data streaming).

6. Compaction interaction

Packs are immutable like all DreamDB Objects. The bucket compactor in 0021 targets the capacity floor K, not F=1, and is not a Fragment repacker.

Dataset::repack_fragments(field, limits) is an explicit, whole-field operation for ordinary Image/Audio/Video Fragment Tracks. It does not support chunked Items, CMAF Tracks or VideoItem Tracks; those refuse before writes. It uses the existing 6-tuple encoding, preserves every entry's anchor interval and Item bytes, and changes only the selected field's active Track in one ordinary Manifest/Ref CAS. Physical Fragment addresses/offsets may change; callers must resolve locators against the selected snapshot. Old snapshots and their content remain readable.

The operator supplies positive max_items, max_input_bytes and max_pack_bytes limits. Planning enumerates the field's Track metadata first; the item limit, summed Item bytes and summed physical input Object sizes are checked before reading payloads or writing. No unbounded-payload or constant-metadata-memory claim is made. A field over the input limits refuses; v1 does not silently repack a prefix. Output groups follow Track order and contain at most Schema pack_items Items and max_pack_bytes bytes (the latter is at most u32::MAX). An individual Item larger than that target remains a single unpacked Fragment; it is neither split nor rejected for exceeding the output target, but still counts against input limits. Packing must be enabled in Schema; this operation does not change Schema.

All reads, slicing and child-Manifest validation precede writes. Repacking publishes only if the candidate has strictly fewer distinct current-root Fragment addresses. Otherwise it is a zero-write no-op, including repeated execution on its own output. Different legal packing geometry alone does not justify rewriting. Normal stale-parent/CAS rules apply; no automatic retry retargets a newer snapshot. No GC is invoked: retained parents, other Refs and pinned snapshots can keep all old Objects alive. Reduced current-root count is not reclaimed storage.

7. Why this design (vs alternatives)

AlternativeWhy rejected
Pack header CBOR with per-item index inside the packDuplicates the Track Object's index; two sources of truth; format-evolution problem.
Separate pack/ path prefix at the address layerAdds visible path complexity; consumer needs to know two paths exist. Content-addressing already disambiguates.
New InlineObjectIndex::FragmentPack variant in Track ObjectAdds a new variant to 129+ match sites across the codebase. Not justified for what is a small per-entry shape change.
Pack with internal byte-range index AND Track-level pack_offsetStrictly more bytes for the same information.
Chunking (0014) for packingChunking splits ONE big item into N PUTs. Packing bundles N small items into ONE PUT. Opposite directions; can't compose.

8. Conformance requirements (catalogued in 0009)

The conformance categories for fragment packs are catalogued at 0009 §8.8:

  1. Round-trip a pack: write N=8 items with pack_items=4 → expect 2 packs in S3 → read back, byte-equality vs originals.
  2. Mixed packs in one Track: append in 3 batches with different N each batch → expect 3 separate packs, item count matches.
  3. Mixed packed and unpacked Fields in one Schema: confirm per-Field independence of the pack_items knob.
  4. Reject pack + chunk combination: ingest a sample with both chunk_size and pack_items > 1 set on the same field → writer MUST refuse, OR writer MUST silently degrade to chunking (current behavior).
  5. Byte-range fetch: a single-item GET with pack_offset set returns exactly [pack_offset .. pack_offset + byte_size] from the pack.

9. Measured impact

10K imagenet S3 ingest, laptop → us-east-1, fair A/B (back-to-back warm runs, 2026-05-22):

ModeS3 image objectsWall-clockRate
Unpacked (1 PUT/item)10,00072.0 s138.9 /s
Packed (pack_items=32)31366.1 s151.3 /s

The packing reduces S3 PUT count 32×. Wall-clock improves ~8% — modest because CLIP-MPS at ~150 /s is the dominant cost on this hardware. On GPU hardware where CLIP encoding is no longer the floor, the network savings translate to a much larger wall-clock win (estimated 5–10× per design/0008 §"Expected payoff").

The structural change is required for billion-scale ingest where individual-item PUTs run into S3 request-rate limits and dollar cost. With packing, network cost scales as O(N / pack_items) rather than O(N).

10. Open questions

The #345 public-boundary workload (64 distinct 1 KiB Items, pack_items=8, one 64-item scan batch, in-memory backend) measured:

Append batchCurrent-root packsPayload bytesPack bytes min/maxScan pack GETs
1 Item64655361024 / 102464
64 Items8655368192 / 81928

Counts exclude metadata and retained historical roots, and distinct payloads exclude content deduplication as the explanation. Explicit repacking with an 8192-byte output target reduces the first case to 8 current-root packs, preserving both new and historical reads. This is request-count evidence, not a WAN timing, cloud-cost, universal byte optimum or physical-storage-reclamation measurement.

OQDescription
OQ-97RESOLVED — schema-pinned. pack_items is a persisted Schema field, declared per field at schema-construction time and serialised into the Schema Object (dreamdb-dataset schema.rs; exposed that way by the Python bindings' field constructors). append_many takes no packing parameter and no per-call override exists, so the packing geometry of a field is a property of the dataset rather than of whoever happens to be writing to it.
OQ-98RESOLVED — workload/operator policy, not a universal optimum. Schema pack_items remains the count ceiling; explicit repacking also accepts a bounded caller byte target (§6). The measurements above justify neither a default 2 MB target nor a universal performance optimum.
OQ-99RESOLVED — explicit bounded repacking, §6, after observed current-root pack/request sprawl. No automatic execution, GC or historical rewrite.

Future changes to default sizing require their own workload evidence.