Skip to main content
CyborgDB supports BM25 full-text search on designated metadata fields, alongside vector search. This unlocks two new query paths on top of the existing vector query: Text-search data is stored encrypted: the backing store never sees a readable term, which terms a document contains, or how many documents contain a given term.

How it works

  • BM25 is enabled by designating fields, not by a flag. An index with at least one full_text field supports text search. An index with none is unaffected.
  • Text is sourced from metadata field values, not from contents. contents stays arbitrary binary and is never analyzed.
  • English only (currently): lowercase → split on non-alphanumeric → stopword filter → Porter2 stem.
  • Designations are fixed for the index’s lifetime. They cannot currently be changed after creation.
  • Fusion is by rank, not score. BM25 relevance and vector distance have no common unit, so hybrid queries fuse positions: score(d) = Σ w_leg / (k + rank_leg(d)). alpha sets the leg weights, rrf_k damps how much a single leg’s top hit dominates, and window_mult sets how deep each leg retrieves before fusing.

Designating fields

At index creation, mark fields as full_text in metadata_schema — or use the text_fields shorthand:
The longhand is equivalent, and is what you need when a field wants other policy too:
full_text=True requires filterable=False, and is incompatible with pattern=True. In Python, {"full_text": True} alone implies filterable: False — only writing filterable: True yourself is a contradiction. If a field needs to be both searchable and exact-match filterable, store it twice under two names.

Ingest

Nothing special is needed at upsert — text comes from the designated fields’ metadata values:
A full_text field that is missing or whose value is not a JSON string is skipped for that item (the value is never logged) and excluded from that field’s document count. An empty or all-stopword value is legal and simply contributes no terms.

Full-text query — no vector

Rows come back descending by BM25 score (ties by id). score is present only when there is one — a filter-only query_metadata() call has nothing to score, so the key is absent rather than None.

Narrowing the searched fields, weighting, and AND

Composing with a filter

A filter combines with text as a pre-filter — it resolves first, then only its survivors are scored. Document frequency and per-field corpus stats stay corpus-global, so a document’s score never depends on the filter:
order_by is not supported together with text — text results rank by score.

Hybrid query — BM25 + vector

Set text=... on query() to fuse BM25 with vector similarity:
Hybrid results carry score (fused relevance, larger = better) and never distance — a document matched by text alone has no vector distance to report. A query vector is required even at alpha=0 — it fixes the batch shape. For text search with no vector, use query_metadata(text=...).
Hybrid queries narrow which filters are accepted. $regex on a non-pattern field, and any predicate on an explicitly non-filterable field, work in a pure query() but are rejected once text is set, because the text leg can only be filtered on fields resolvable from the encrypted metadata index.

Tuning

Two non-obvious properties worth knowing before tuning:
  • No positive rrf_k lets one leg’s rank-1 outrank two legs’ rank-2. Solving 1/(k+1) = 2/(k+2) gives k = 0. If you want a single leg to dominate, use alpha, not rrf_k.
  • rrf_k is only meaningful as deep as the legs actually retrieved. At window_mult = 3 and top_k = 10 each leg is 30 deep, and every value of k above ~28 makes the same decision — so sweeping rrf_k alone at a shallow window will look like the knob does nothing. Sweep the two together.

Constraints and gotchas

  • full_text + filterable=True (explicit) is an error, and full_text + pattern=True is an error always — a field can’t be both full-text and pattern-indexed. If you need a field both searchable and exact-match filterable, store it twice under two names.
  • Designations are fixed for the index’s lifetime. They cannot currently be changed after creation.
  • order_by cannot be combined with text. Text results rank by score.
  • Not currently supported: phrase queries, NOT/exclusion, nested boolean expressions, non-English analysis, and dual-indexing a field as both analyzed and exact-match.
  • The analyzer is versioned. The pipeline version, stopword-list hash, and stemmer identity are stamped into the index at creation. Opening an index with a build whose analyzer differs is a loud error, not silent recall loss — every existing document was indexed by the old analyzer.
  • Multi-tenant guidance: use index-per-tenant. Corpus statistics are per index, so co-mingling tenants behind a tenant_id filter gives every tenant blended term rarity — BM25 statistics are deliberately global and never scoped to a filtered subset.

Common errors