Semantic search
DreamDB searches vectors; it does not make them. Whatever produced the embeddings at ingest time has to produce the query vector too — the same model, the same normalization.
The rule that matters
A query encoded by one model scored against an index built by another computes fine, returns plausible-looking scores, and is confidently wrong. Nothing in the data catches it. Spec 0024 exists because of this.
So: if a dataset's embeddings came from CLIP ViT-B/32 at 512 dimensions, the query vector must come from CLIP ViT-B/32 at 512 dimensions. Check what you have before writing any UI:
Encoding the query in the page
transformers.js runs the CLIP text tower in the browser. The model downloads once and is cached by the browser:
normalize: true matters: the cosine algorithms assume unit-length vectors.
Querying
That is the whole search path. No server, no vector database process — the index lives in the address space, so the SDK fetches the buckets a query lands in and scores them locally.
Showing the results
Hits are anchors and scores. Join them to whatever you want to display, exactly as in the React example:
When results come back empty
On a small dataset this is usually not a bug. Partitioning algorithms send a query to the bucket its direction falls in, and with a few hundred records most buckets are empty — a random query vector legitimately finds nothing. Two responses:
- Raise
probeCountto widen the search across neighbouring buckets:queryVector(field, vec, { topK: 24, probeCount: 16 }). - Confirm the path works by querying with a vector you know is in the set. It should come back first, at a score of 1.000.
Results that are ordered plausibly but wrong are a different failure, and the usual cause is the one at the top of this page: the query vector came from a different model, or a different normalization, than the indexed vectors. Check track.modality against what encoded your query before looking anywhere else.
Text search without embeddings
If the field is text rather than vectors, queryText runs BM25 with a built-in tokenizer, and needs no model at all:
queryHybrid fuses both — a lexical sub-query and a dense one — with RRF, linear, max, or Pareto fusion. Both are described in the API reference; building the text index behind them is a Python or CLI job.