Skip to main content
CyborgDB supports BM25 full-text search on designated metadata fields, alongside vector search. This unlocks two new query paths on top of the existing vector query: Text-search data is stored encrypted: the backing store never sees a readable term, which terms a document contains, or how many documents contain a given term.

How It Works

  • BM25 is enabled by designating fields, not by a flag. An index with at least one full_text field supports text search. An index with none has no full-text overhead.
  • Text is sourced from metadata field values, not from contents. contents stays arbitrary binary and is never analyzed.
  • English only (currently): lowercase → split on non-alphanumeric → stopword filter → Porter2 stem.
  • Designations are fixed for the index’s lifetime. They currently cannot be changed after creation.
  • Fusion is by rank, not score. BM25 relevance and vector distance have no common unit, so hybrid queries fuse positions: score(d) = Σ w_leg / (k + rank_leg(d)). alpha sets the leg weights, rrf_k damps how much a single leg’s top hit dominates, and window_mult sets how deep each leg retrieves before fusing.

Designating Fields

At index creation, mark fields as full_text in metadata_schema — or use the text_fields shorthand:
The longhand — metadata_schema={"title": {"full_text": true}} — is equivalent, and is what you need when other fields want policy too (e.g. a filterable or pattern field alongside). The BM25 scoring constants can also be tuned at creation via bm25_k1 (term-frequency saturation, default 1.2) and bm25_b (length normalization, default 0.75); both require at least one full_text field.
full_text: true implies filterable: false and is incompatible with pattern: true — a field is either analyzed for BM25 or exact-match indexed, never both. Writing an explicit filterable: true alongside full_text: true is rejected. If a field needs to be both text-searchable and exact-match filterable, store it twice under two names.

Ingesting Text

Nothing special is needed at upsert — text comes from the designated fields’ metadata values:
A full_text field’s value must be a string. Items where the field is missing or non-string (including a list of strings) are skipped for that field — the upsert still succeeds. An empty or all-stopword value is legal and simply contributes no terms.

Full-Text Query — No Vector

Pass text to query_metadata() to run a pure BM25 search. Rows come back as {"id", "score"}, descending by BM25 score:
score is present only when there is one — a filter-only query_metadata() call has nothing to score, so the key is absent rather than null. Three optional knobs shape the text search:
  • text_fields — search only some of the designated fields (naming a non-full_text field is rejected).
  • text_field_weights — per-field weights, parallel to the searched fields (name them explicitly when weighting).
  • require_all_terms — require every term to match (AND) instead of any (OR, the default).
A filters argument given alongside text acts as a pre-filter — it resolves first, then only its survivors are scored. Document frequency and per-field corpus stats stay corpus-global, so a document’s score never depends on the filter. order_by is not supported together with text — text results rank by score.

Hybrid Query — BM25 + Vector

Set text on query() to fuse BM25 with vector similarity:
Hybrid results carry score (fused relevance, larger = better) and never distance — a document matched by text alone has no vector distance to report. A query vector is required even at alpha=0 — it fixes the batch shape. For text search with no vector, use query_metadata() with text.

Three-Way Query — Vector + BM25 + Metadata

Adding filters on top of a hybrid query composes all three legs in one request. The filter is resolved once, before retrieval, and the surviving IDs constrain both the vector and BM25 legs — so both rankings are computed over the same filtered candidate set, then fused by RRF. A filter that matches nothing short-circuits to an empty result without running either leg:
alpha still applies: 0 reduces the query to BM25 + metadata, 1 to vector + metadata, and either exact endpoint skips the dead leg’s retrieval entirely. For BM25 + metadata with no vector at all, use query_metadata() with text and filters instead.
Any query with text set accepts only filters the metadata index can resolve exactly. $regex on a non-pattern field, and any predicate on an explicitly non-filterable field, work in a pure query() but are rejected once text is set — otherwise those filters could not be applied to the full-text results, and they would be silently returned unfiltered. In the example above, author works because it is an ordinary filterable field, not a full_text one.

Tuning

Two non-obvious properties worth knowing before tuning:
  • No positive rrf_k lets one leg’s rank-1 outrank two legs’ rank-2. If you want a single leg to dominate, use alpha, not rrf_k. At alpha=0 or alpha=1 the dead leg’s retrieval is skipped entirely.
  • rrf_k is only meaningful as deep as the legs actually retrieved. At window_mult=3 and top_k=10 each leg is 30 deep, and every value of rrf_k above ~28 makes the same decision — sweeping rrf_k alone at a shallow window will look like the knob does nothing. Sweep the two together.

Reading Back the Config

The recorded BM25 config — k1, b, and the analyzer_version the corpus was indexed with — is reported on the index handle, or null when the index has no full_text field:

Constraints and Gotchas

  • full_text + explicit filterable: true is an error, and full_text + pattern: true is an error always. If you need a field both text-searchable and exact-match filterable, store it twice under two names.
  • Designations are fixed for the index’s lifetime. They currently cannot be changed after creation.
  • order_by cannot be combined with text. Text results rank by score.
  • Only single string values are analyzed. A full_text field holding a number, boolean, object, or a list of strings is skipped for that item — join multi-part text into one string before upserting.
  • Not currently supported: phrase queries, NOT/exclusion, nested boolean expressions, non-English analysis, and dual-indexing a field as both analyzed and exact-match.
  • The analyzer is versioned. The pipeline version, stopword-list hash, and stemmer identity are stamped into the index at creation (see analyzer_version above). Opening an index with a build whose analyzer differs is a loud error, not silent recall loss.
  • Multi-tenant guidance: use index-per-tenant. Corpus statistics are per index, so co-mingling tenants behind a tenant_id filter gives every tenant blended term rarity — BM25 statistics are deliberately global and never scoped to a filtered subset.

Common Errors

Create-time schema violations are rejected by request validation (HTTP 422); query-time and cross-parameter violations are enforced by the index core (HTTP 400). The SDKs surface both as errors/exceptions with the reason.

API Reference

For more information on full-text and hybrid search, refer to the API reference:

REST API: Query

Hybrid query on /v1/vectors/query

REST API: Query Metadata

Full-text query on /v1/vectors/query_metadata

Python SDK Reference

API reference for query() and query_metadata() in Python

JS/TS SDK Reference

API reference for query() and queryMetadata() in JavaScript/TypeScript

Go SDK Reference

API reference for Query() and QueryMetadata() in Go

Create Index (REST API)

Designating full_text fields at creation