DreamDB

Quickstart

Version scope: Python 0.0.12 / JavaScript 0.5.4. Entity keys, progressive geometry, structured ragged/CSR and additional index capabilities are included in these releases. See feature examples and limits for the supported boundaries; older packages may not expose these methods.

Reading takes three lines and no credentials. Writing takes a backend that can PUT. Both are below, and both were run against the published package.

Read and query

Point Space.fromUri at a ref. Passing null for the backend uses plain fetch against the URI's base, which is all a readable bucket needs:

js
import { Space } from "@dreamlake/dreamdb";

const space = await Space.fromUri("http://localhost:8791/refs/photos", null);

console.log(space.manifestHash);            // d2wn2nmxvqlmt7j6cvvd3r2x…
console.log(space.tracks().map(t => t.key)); // [ 'embedding', 'label' ]

A vector query takes the field name, a Float32Array, and options:

js
const query = new Float32Array(64).map(() => Math.random() - 0.5);
const hits = await space.queryVector("embedding", query, { topK: 3 });
// [ { anchor: 1785307993507656000, score: 0.381 },
//   { anchor: 1785307993507526000, score: 0.256 },
//   { anchor: 1785307993507639000, score: 0.255 } ]

Anchors identify records. To turn them into values, read a scalar column — it comes back as a Map keyed by the anchor as a string:

js
const labels = await space.readScalarColumn(space.trackFor("label"), null);
labels.size;                       // 300
labels.get(String(hits[0].anchor)); // 'cat'

Or go the other way, from a value to its anchors:

js
const cats = await space.queryScalar("label", "eq", "cat"); // 100 anchors

Don't have a dataset yet? The Quick Start writes one to a local directory with Python in about twenty lines, then serves it over HTTP — which is exactly the http://localhost:8791 above.

Write

Writing needs a backend object with put. The bundled S3Backend is read-only, and the write entry points say so rather than failing later:

Error: Authoring.create: backend is missing put — it looks like a read-only (v1)
backend. Contract v2 needs get(path, range?) -> {bytes, etag}, put(path, bytes,
{ifMatch, ifNoneMatchStar}), and head(path).

For a server-side script, a fetch wrapper is enough. This one talks to MinIO:

js
const BASE = "http://localhost:9000/demo";

export default {
  async get(path, range) {
    // `range` is half-open [start, end); HTTP Range is inclusive, hence end - 1.
    const headers = range ? { Range: `bytes=${range.start}-${range.end - 1}` } : {};
    const r = await fetch(`${BASE}/${path}`, { headers });
    if (!r.ok && r.status !== 206) throw new Error(`get ${path}: ${r.status}`);
    return { bytes: new Uint8Array(await r.arrayBuffer()),
             etag: r.headers.get("etag") ?? undefined };
  },
  async head(path) {
    const r = await fetch(`${BASE}/${path}`, { method: "HEAD" });
    return r.ok
      ? { exists: true, etag: r.headers.get("etag") ?? undefined,
          size: Number(r.headers.get("content-length") ?? 0) }
      : { exists: false };
  },
  async put(path, bytes, opts = {}) {
    const headers = {};
    if (opts.ifMatch) headers["If-Match"] = opts.ifMatch;
    if (opts.ifNoneMatchStar) headers["If-None-Match"] = "*";
    const r = await fetch(`${BASE}/${path}`, { method: "PUT", headers, body: bytes });
    if (r.status === 412) return { status: opts.ifMatch ? "casFailed" : "exists" };
    if (!r.ok) throw new Error(`put ${path}: ${r.status}`);
    return { status: "created", etag: r.headers.get("etag") ?? undefined };
  },
};

Then create a dataset and append to it. Fields are declared as plain objects, and each sample names its field types:

js
import { Authoring } from "@dreamlake/dreamdb";
import backend from "./backend.mjs";

const BASE = "http://localhost:9000/demo";

const authoring = await Authoring.create("js-notes", [
  { name: "vec", kind: "embedding", dim: 16, algorithm: "dreamdb.lsh-cosine" },
  { name: "tag", kind: "scalar", valueType: "categorical", required: false },
], BASE, backend);

const writer = authoring.writer();
const t0 = BigInt(Date.now()) * 1_000_000n;

const n = await writer.appendMany(
  Array.from({ length: 40 }, (_, i) => ({
    anchor: t0 + BigInt(i * 1000),
    vec: {
      kind: "embedding",
      algorithm: "dreamdb.lsh-cosine",
      vector: new Float32Array(16).map(() => Math.random() - 0.5),
    },
    tag: { kind: "categorical", value: i % 2 ? "a" : "b" },
  })),
);

console.log("appended", n);                       // appended 40
console.log("commit", await writer.commit());     // dzfu35vaojsiquz4m3km…
console.log("snapshot", await writer.snapshot("v1"));

Nothing is visible to a reader until commit(). That call is what advances the ref, with a compare-and-swap against its current ETag.

To append to a dataset that already exists, skip Authoring entirely:

js
import { Writer } from "@dreamlake/dreamdb";

const writer = await Writer.open(`${BASE}/refs/js-notes`, backend);

Writer is in the browser build too. Authoring is not — see Overview.

Next