DreamDB

Storage backends

A backend is one string. Change it and the same code writes to a local directory, a MinIO container, or a production bucket. There is no server component to run either way.

Backend stringWhat it isPythonJavaScript
file:///abs/pathA directory on diskRead + write—
memory://In-process, discarded on exit. For testsRead + write—
http://host/bucketAny S3-compatible endpointRead + writeRead, and write with a put backend
https://host/bucketThe same, over TLSRead + writeRead, and write with a put backend

JavaScript always goes over HTTP. It has no filesystem path, so a dataset written to file:// has to be served before a browser or Node process can read it — and a plain static server can serve reads but not accept writes.

Local disk

The fastest way to work. Nothing to install, nothing to start:

python
ds = db.Dataset.create("photos", schema, backend="file:///tmp/dreamdb-quickstart")

To read that dataset from JavaScript, serve the directory:

bash
npx http-server /tmp/dreamdb-quickstart -p 8791 --cors

--cors matters only for browsers; a Node script can read a server without it.

MinIO

An S3-compatible server in a container. Use it when you want to exercise the real HTTP path locally, or to share a dataset across machines on your network:

bash
docker run -d --name dreamdb-minio \
  -p 9000:9000 -p 9001:9001 \
  -e MINIO_ROOT_USER=dreamdb \
  -e MINIO_ROOT_PASSWORD=dreamdbsecret \
  minio/minio:latest server /data --console-address ":9001"

docker exec dreamdb-minio mc alias set local http://localhost:9000 dreamdb dreamdbsecret
docker exec dreamdb-minio mc mb local/demo
docker exec dreamdb-minio mc anonymous set public local/demo

The backend string is then http://localhost:9000/demo. Making the bucket public keeps development simple; it also means anyone who can reach the port can read it.

S3, R2, B2, Wasabi

Any store that speaks S3-style GET and PUT works. Point at the bucket:

python
ds = db.Dataset.create("photos", schema,
                       backend="https://my-bucket.s3.us-east-1.amazonaws.com")

Writes are signed with SigV4, picked up from the standard environment variables:

bash
export AWS_ACCESS_KEY_ID=…
export AWS_SECRET_ACCESS_KEY=…
export AWS_REGION=us-east-1

No code changes are needed to move from MinIO to production. Set the variables and change the URL.

Backends in JavaScript

JavaScript takes a backend object rather than a URL string. Only get is required, and a get-only backend is a read-only backend — the write entry points check for put up front and refuse by name:

ts
interface Backend {
  get(path: string, range?: { start: number; end: number })
    : Promise<Uint8Array | { bytes: Uint8Array; etag?: string; totalLength?: number }>
  getStream?(path: string)
    : Promise<ReadableStream<Uint8Array> | {
        stream: ReadableStream<Uint8Array>;
        etag?: string;
        totalLength?: number;
      }>
  head?(path: string): Promise<{ exists?: boolean; etag?: string; size?: number }>
  put?(path: string, bytes: Uint8Array,
       opts?: { ifMatch?: string; ifNoneMatchStar?: boolean })
    : Promise<{ status: "created" | "exists" | "casFailed"; etag?: string }>
  delete?(path: string): Promise<void>
  list?(prefix: string): Promise<string[]>
}

Passing null instead uses plain fetch against the URI's base, which is enough to read a public bucket:

range is half-open — [start, end), the shape the Rust core uses. An HTTP Range header is inclusive, so a fetch-based backend must send bytes=${start}-${end - 1}. Sending end fetches one byte too many.

getStream is optional. Existing backends remain readable through get; implementing a bounded stream lets DreamDB scan historical oversized inline Tracks without first aggregating the complete object in wasm memory. A structured get or getStream response should include totalLength when known so range verification can distinguish the returned slice from the complete object.

js
const space = await Space.fromUri("https://bucket.s3.amazonaws.com/refs/photos", null);

The bundled S3Backend is read-only. For server-side writes, a fetch wrapper of about twenty lines is enough — get with a Range header, head, and put translating ifMatch to If-Match and ifNoneMatchStar to If-None-Match: *, returning casFailed on a 412.

Writing from a browser

Credentials must not reach the page. PresignedBackend keeps them out: your server mints short-lived presigned PUT URLs, and performs the ref's compare-and-swap itself.

js
import { Writer, PresignedBackend } from "@dreamlake/dreamdb";

const backend = new PresignedBackend({
  readBase: "https://bucket.s3.amazonaws.com",
  mintPut: async (paths) => (await postJson("/api/dreamdb/sign", { paths })).urls,
  commitRef: async (path, opts, bytes) =>
    await postJson("/api/dreamdb/ref", { path, bytesBase64: base64(bytes), ...opts }),
});

const w = await Writer.open("https://bucket.s3.amazonaws.com/refs/photos", backend);

Anchors are nanoseconds, and in JavaScript you pass a bigint. A number is accepted only below Number.MAX_SAFE_INTEGER; above that the SDK throws instead of rounding.

Two rules for anything in front of a bucket

Range should be honoured. DreamDB reads exact vectors by byte offset, and asks for half-open [start, end) ranges that the HTTP layer sends as inclusive bytes=start-(end-1). A proxy or server that ignores Range and replies 200 with the whole object does not break correctness — the connectors detect it and slice locally — but every ranged read then transfers the whole object, which makes large datasets impractical. Such a backend also fails Range conformance and cannot claim the corresponding performance profile.

CORS must expose ETag. Refs advance by compare-and-swap, so the SDK needs to read the ref's current ETag. Without ExposeHeaders, the browser hides that header, the SDK falls back to create-only, and the second commit to any ref fails while the first succeeded:

json
{
  "CORSRules": [
    {
      "AllowedOrigins": ["https://your-app.example"],
      "AllowedMethods": ["GET", "HEAD", "PUT"],
      "AllowedHeaders": ["*"],
      "ExposeHeaders": ["ETag", "Content-Range", "Content-Length"],
      "MaxAgeSeconds": 3600
    }
  ]
}

Layout on disk

Whatever the backend, the object layout is the same, which is why a directory served over HTTP is indistinguishable from a bucket:

<root>/
├── refs/<name>            # the pointer a ref name resolves to
├── manifests/<hash>       # immutable versions, each naming its parent
├── genesis                # the timeline root
├── spatial-index/         # trained indexes
├── scalar-index/
└── <timeline>/…           # track data, addressed by content

Nothing here is a database file. Every object is content-addressed and immutable except refs/<name>, which is a single pointer updated with compare-and-swap.