# Storage backends

A backend is one string. Change it and the same code writes to a local directory, a MinIO container, or a production bucket. There is no server component to run either way.

| Backend string | What it is | Python | JavaScript |
|---|---|---|---|
| `file:///abs/path` | A directory on disk | Read + write | — |
| `memory://` | In-process, discarded on exit. For tests | Read + write | — |
| `http://host/bucket` | Any S3-compatible endpoint | Read + write | Read, and write with a `put` backend |
| `https://host/bucket` | The same, over TLS | Read + write | Read, and write with a `put` backend |

JavaScript always goes over HTTP. It has no filesystem path, so a dataset written to `file://` has to be served before a browser or Node process can read it — and a plain static server can serve reads but not accept writes.

## Local disk

The fastest way to work. Nothing to install, nothing to start:

```python
ds = db.Dataset.create("photos", schema, backend="file:///tmp/dreamdb-quickstart")
```

To read that dataset from JavaScript, serve the directory:

```bash
npx http-server /tmp/dreamdb-quickstart -p 8791 --cors
```

`--cors` matters only for browsers; a Node script can read a server without it.

## MinIO

An S3-compatible server in a container. Use it when you want to exercise the real HTTP path locally, or to share a dataset across machines on your network:

```bash
docker run -d --name dreamdb-minio \
  -p 9000:9000 -p 9001:9001 \
  -e MINIO_ROOT_USER=dreamdb \
  -e MINIO_ROOT_PASSWORD=dreamdbsecret \
  minio/minio:latest server /data --console-address ":9001"

docker exec dreamdb-minio mc alias set local http://localhost:9000 dreamdb dreamdbsecret
docker exec dreamdb-minio mc mb local/demo
docker exec dreamdb-minio mc anonymous set public local/demo
```

The backend string is then `http://localhost:9000/demo`. Making the bucket public keeps development simple; it also means anyone who can reach the port can read it.

## S3, R2, B2, Wasabi

Any store that speaks S3-style `GET` and `PUT` works. Point at the bucket:

```python
ds = db.Dataset.create("photos", schema,
                       backend="https://my-bucket.s3.us-east-1.amazonaws.com")
```

Writes are signed with SigV4, picked up from the standard environment variables:

```bash
export AWS_ACCESS_KEY_ID=…
export AWS_SECRET_ACCESS_KEY=…
export AWS_REGION=us-east-1
```

No code changes are needed to move from MinIO to production. Set the variables and change the URL.

## Backends in JavaScript

JavaScript takes a backend object rather than a URL string. Only `get` is required, and a `get`-only backend is a read-only backend — the write entry points check for `put` up front and refuse by name:

```ts
interface Backend {
  get(path: string, range?: { start: number; end: number })
    : Promise<Uint8Array | { bytes: Uint8Array; etag?: string; totalLength?: number }>
  getStream?(path: string)
    : Promise<ReadableStream<Uint8Array> | {
        stream: ReadableStream<Uint8Array>;
        etag?: string;
        totalLength?: number;
      }>
  head?(path: string): Promise<{ exists?: boolean; etag?: string; size?: number }>
  put?(path: string, bytes: Uint8Array,
       opts?: { ifMatch?: string; ifNoneMatchStar?: boolean })
    : Promise<{ status: "created" | "exists" | "casFailed"; etag?: string }>
  delete?(path: string): Promise<void>
  list?(prefix: string): Promise<string[]>
}
```

Passing `null` instead uses plain `fetch` against the URI's base, which is enough to read a public bucket:

`range` is **half-open** — `[start, end)`, the shape the Rust core uses. An HTTP `Range` header is inclusive, so a fetch-based backend must send `bytes=${start}-${end - 1}`. Sending `end` fetches one byte too many.

`getStream` is optional. Existing backends remain readable through `get`; implementing a bounded stream lets DreamDB scan historical oversized inline Tracks without first aggregating the complete object in wasm memory. A structured `get` or `getStream` response should include `totalLength` when known so range verification can distinguish the returned slice from the complete object.

```js
const space = await Space.fromUri("https://bucket.s3.amazonaws.com/refs/photos", null);
```

The bundled `S3Backend` is read-only. For server-side writes, a `fetch` wrapper of about twenty lines is enough — `get` with a `Range` header, `head`, and `put` translating `ifMatch` to `If-Match` and `ifNoneMatchStar` to `If-None-Match: *`, returning `casFailed` on a 412.

### Writing from a browser

Credentials must not reach the page. `PresignedBackend` keeps them out: your server mints short-lived presigned `PUT` URLs, and performs the ref's compare-and-swap itself.

```js
import { Writer, PresignedBackend } from "@dreamlake/dreamdb";

const backend = new PresignedBackend({
  readBase: "https://bucket.s3.amazonaws.com",
  mintPut: async (paths) => (await postJson("/api/dreamdb/sign", { paths })).urls,
  commitRef: async (path, opts, bytes) =>
    await postJson("/api/dreamdb/ref", { path, bytesBase64: base64(bytes), ...opts }),
});

const w = await Writer.open("https://bucket.s3.amazonaws.com/refs/photos", backend);
```

Anchors are nanoseconds, and in JavaScript you pass a `bigint`. A `number` is accepted only below `Number.MAX_SAFE_INTEGER`; above that the SDK throws instead of rounding.

## Two rules for anything in front of a bucket

**`Range` should be honoured.** DreamDB reads exact vectors by byte offset, and asks for half-open `[start, end)` ranges that the HTTP layer sends as inclusive `bytes=start-(end-1)`. A proxy or server that ignores `Range` and replies 200 with the whole object does not break correctness — the connectors detect it and slice locally — but every ranged read then transfers the whole object, which makes large datasets impractical. Such a backend also fails Range conformance and cannot claim the corresponding performance profile.

**CORS must expose `ETag`.** Refs advance by compare-and-swap, so the SDK needs to read the ref's current ETag. Without `ExposeHeaders`, the browser hides that header, the SDK falls back to create-only, and the *second* commit to any ref fails while the first succeeded:

```json
{
  "CORSRules": [
    {
      "AllowedOrigins": ["https://your-app.example"],
      "AllowedMethods": ["GET", "HEAD", "PUT"],
      "AllowedHeaders": ["*"],
      "ExposeHeaders": ["ETag", "Content-Range", "Content-Length"],
      "MaxAgeSeconds": 3600
    }
  ]
}
```

## Layout on disk

Whatever the backend, the object layout is the same, which is why a directory served over HTTP is indistinguishable from a bucket:

```
<root>/
├── refs/<name>            # the pointer a ref name resolves to
├── manifests/<hash>       # immutable versions, each naming its parent
├── genesis                # the timeline root
├── spatial-index/         # trained indexes
├── scalar-index/
└── <timeline>/…           # track data, addressed by content
```

Nothing here is a database file. Every object is content-addressed and immutable except `refs/<name>`, which is a single pointer updated with compare-and-swap.
