DreamDB

DreamDB Specification — 0023: Performance Profiles

Status: Draft. Companion to 0009 (Conformance). Defines the form of a DreamDB performance profile and the boundary between performance and conformance. Builds on 0004, 0007, 0010.


1. Purpose

The DreamDB spec contains performance narrative — latency budgets, recall figures, bytes-fetched estimates, sizing tables (0004 §7.3/§8, 0007, 0010). These guide implementers and justify design choices, but they are not protocol requirements: whether a query returns in 100 ms or 500 ms depends on the workload (corpus size, dimensionality, vector distribution) and the environment (storage round-trip, bandwidth, CPU/GPU) — neither of which the protocol controls.

Conflating those numbers with correctness leaves an implementer unable to tell whether missing 100 ms means non-conformant or merely slower. This document removes that ambiguity by defining performance as a separate, named, parameterized artifact — a performance profile — that an implementation claims, distinct from the pass/fail conformance Tier of 0009.

This document is not a benchmark suite and pins no minimum performance. It defines (a) the conformance/performance boundary, (b) the structure of a profile, and (c) how implementations report which profiles they meet. One worked profile is given as a template (§6); it is illustrative, not mandatory.

2. The conformance / performance boundary (normative)

Per 0009 §4.4, a test criterion is one of two classes:

  • Conformance criterion — deterministic and environment-independent (byte-identical, functional-equivalence, behavioral-assertion). Gates conformance. For example, a snapshot query issues no LISTs. An exact GET count needs a fixture fixing Bucket splits, cache state, coalescing and retries; table count alone does not determine it (0009 §4.4).
  • Performance-profile target — depends on workload and/or environment (tolerance-bound: latency, throughput, recall, wall-clock bytes). Advisory. Defined here.

Litmus test. If the same inputs always yield the same pass/fail on any conformant implementation regardless of hardware, it is a conformance criterion. If the answer can change with the machine, the network, or the corpus, it is a performance-profile target.

A profile target is never a conformance gate. An implementation that satisfies the protocol but meets zero performance profiles is fully DreamDB-conformant; it simply claims no profile.

3. Anatomy of a performance profile

A performance profile is a named tuple:

profile := (id, description, workload, environment_class, targets)
  • id — stable identifier, <family>.<version> (e.g. billion-scale-ann.v0). Profiles version independently of the protocol (§8).
  • workload — the dataset and query shape the targets are measured against: item count, modality/dimensionality, vector distribution, index parameters (N, tables, M, compressor), query mix, top_k, recall operating point.
  • environment_class — the assumed deployment: storage round-trip (p50/p99), bandwidth, HTTP version, client CPU/GPU, cache state (cold vs warm). Targets are only comparable within an environment class. In particular, the latency profiles assume HTTP/2 (or HTTP/3) end-to-end for the parallel-ranged-GET hot path (0005 §6.1); an HTTP/1.1 deployment is fully DreamDB-conformant (0005 §6) but falls outside these profiles, since HTTP/1.1 serializes those GETs across a bounded connection pool.
  • targets — the tolerance-bound metrics: e.g. recall@10 ≥ 0.90, hot_query_p95 ≤ 100 ms, cold_query_p95 ≤ 200 ms, bytes_fetched_per_query ≤ 12 MB, ingest_throughput ≥ N items/s.

A target without its workload + environment_class is meaningless; the three travel together.

4. Profile document format

Profiles are exchanged as JSON, parallel to the 0009 §4 vector format:

json
{
  "id": "billion-scale-ann.v0",
  "description": "1B-vector cosine ANN, warm-cache hot path, S3-class storage.",
  "workload": {
    "modality": "embedding.f32.dim=768.bucketed",
    "item_count": 1000000000,
    "distribution": "normalized-random | clustered:<k> | dataset:<name>",
    "index": { "algorithm": "dreamdb.lsh-cosine", "spatial_bits": 18, "tables": 1, "prefix_truncate_M": 14 },
    "query": { "top_k": 10, "recall_operating_point": 0.90, "mix": "uniform-random-queries" }
  },
  "environment_class": {
    "storage": { "kind": "object-store", "get_rtt_p50_ms": 15, "get_rtt_p99_ms": 60, "bandwidth_gbps": 10 },
    "http": "h2",
    "client": { "cpu": "x86-64 16-core", "gpu": "none" },
    "cache": "warm"
  },
  "targets": {
    "recall_at_10": { "min": 0.90 },
    "hot_query_latency_p95_ms": { "max": 100 },
    "cold_query_latency_p95_ms": { "max": 200 },
    "bytes_fetched_per_query_mb": { "max": 12 }
  }
}

A profile report (what an implementation publishes) pairs the profile id with the measured values and the concrete environment they were measured on:

json
{
  "profile": "billion-scale-ann.v0",
  "met": true,
  "measured": { "recall_at_10": 0.93, "hot_query_latency_p95_ms": 78, "cold_query_latency_p95_ms": 165, "bytes_fetched_per_query_mb": 11.4 },
  "environment": { "storage": "AWS S3 us-east-1", "client": "c7i.4xlarge", "http": "h2", "date": "2026-06-25" }
}

5. Reporting

An implementation's published report (0009 §11) carries two independent sections:

  1. A conformance section — the pass/fail Tier and per-category results (the only thing that decides conformance).
  2. A performance section — zero or more profile reports (§4). met: false or an absent profile is informational; it never changes the conformance section.

Profiles are comparable across implementations only when workload and environment_class match. Reporting measured values under a different environment than a profile's environment_class is permitted but MUST state the actual environment; it is then a data point, not a claim that the profile is met.

6. Worked profile (informative)

billion-scale-ann.v0 in §4 is a template with N=18 and the workload and environment declared there. Its 100 ms / 12 MB targets are illustrative, not derived from the different N=22 / PQ scenario in 0004. The sample report is also illustrative, not a recorded benchmark. A real report must pin the actual index/compressor, corpus and query set before comparing measured results to those targets. Meeting no performance profile does not make an implementation protocol-nonconformant.

7. Relationship to conformance (normative)

  • Conformance (0009) and performance (this doc) are orthogonal. Neither implies the other.
  • A conformant implementation MAY meet zero profiles. A non-conformant implementation's profile claims are void.
  • Spec text that states a performance number (latency, throughput, recall figure, bytes-fetched) is informative unless it is also expressible as a 0009 conformance criterion. Where such narrative rests on a normative invariant (e.g. Manifest Supremacy — no list-prefix on the steady-state read path, 0005 §5.3.1), that invariant is stated and tested as a conformance criterion in its own right, independent of any budget.

8. Versioning

Performance profiles version on their own cadence, decoupled from the protocol's semantic version (0009 §12). Re-tuning a target, adding an environment class, or publishing a new workload bumps the profile <version> (e.g. billion-scale-ann.v0 → v1) and never constitutes a protocol change. A protocol revision that does not alter byte formats, grammars, or verb semantics leaves every existing profile valid.

9. Open Questions

  • OQ-PP1 (→ community): a shared, version-pinned reference dataset + query set per profile family, so cross-implementation profile reports are directly comparable.
  • OQ-PP2 (→ v0.1): standard environment-class identifiers (e.g. s3-class, local-nvme, in-region-minio) to reduce free-form environment drift.