Text-search data is stored encrypted: the backing store never sees a readable term, which terms a document contains, or how many documents contain a given term.
How It Works
- BM25 is enabled by designating fields, not by a flag. An index with at least one
full_textfield supports text search. An index with none has no full-text overhead. - Text is sourced from metadata field values, not from
contents.contentsstays arbitrary binary and is never analyzed. - English only (currently): lowercase → split on non-alphanumeric → stopword filter → Porter2 stem.
- Designations are fixed for the index’s lifetime. They currently cannot be changed after creation.
- Fusion is by rank, not score. BM25 relevance and vector distance have no common unit, so hybrid queries fuse positions:
score(d) = Σ w_leg / (k + rank_leg(d)).alphasets the leg weights,rrf_kdamps how much a single leg’s top hit dominates, andwindow_multsets how deep each leg retrieves before fusing.
Designating Fields
At index creation, mark fields asfull_text in metadata_schema — or use the text_fields shorthand:
metadata_schema={"title": {"full_text": true}} — is equivalent, and is what you need when other fields want policy too (e.g. a filterable or pattern field alongside). The BM25 scoring constants can also be tuned at creation via bm25_k1 (term-frequency saturation, default 1.2) and bm25_b (length normalization, default 0.75); both require at least one full_text field.
full_text: true implies filterable: false and is incompatible with pattern: true — a field is either analyzed for BM25 or exact-match indexed, never both. Writing an explicit filterable: true alongside full_text: true is rejected. If a field needs to be both text-searchable and exact-match filterable, store it twice under two names.Ingesting Text
Nothing special is needed at upsert — text comes from the designated fields’ metadata values:full_text field’s value must be a string. Items where the field is missing or non-string (including a list of strings) are skipped for that field — the upsert still succeeds. An empty or all-stopword value is legal and simply contributes no terms.
Full-Text Query — No Vector
Passtext to query_metadata() to run a pure BM25 search. Rows come back as {"id", "score"}, descending by BM25 score:
score is present only when there is one — a filter-only query_metadata() call has nothing to score, so the key is absent rather than null.
Three optional knobs shape the text search:
text_fields— search only some of the designated fields (naming a non-full_textfield is rejected).text_field_weights— per-field weights, parallel to the searched fields (name them explicitly when weighting).require_all_terms— require every term to match (AND) instead of any (OR, the default).
filters argument given alongside text acts as a pre-filter — it resolves first, then only its survivors are scored. Document frequency and per-field corpus stats stay corpus-global, so a document’s score never depends on the filter. order_by is not supported together with text — text results rank by score.
Hybrid Query — BM25 + Vector
Settext on query() to fuse BM25 with vector similarity:
score (fused relevance, larger = better) and never distance — a document matched by text alone has no vector distance to report.
A query vector is required even at alpha=0 — it fixes the batch shape. For text search with no vector, use query_metadata() with text.
Three-Way Query — Vector + BM25 + Metadata
Addingfilters on top of a hybrid query composes all three legs in one request. The filter is resolved once, before retrieval, and the surviving IDs constrain both the vector and BM25 legs — so both rankings are computed over the same filtered candidate set, then fused by RRF. A filter that matches nothing short-circuits to an empty result without running either leg:
alpha still applies: 0 reduces the query to BM25 + metadata, 1 to vector + metadata, and either exact endpoint skips the dead leg’s retrieval entirely. For BM25 + metadata with no vector at all, use query_metadata() with text and filters instead.
Any query with
text set accepts only filters the metadata index can resolve exactly. $regex on a non-pattern field, and any predicate on an explicitly non-filterable field, work in a pure query() but are rejected once text is set — otherwise those filters could not be applied to the full-text results, and they would be silently returned unfiltered. In the example above, author works because it is an ordinary filterable field, not a full_text one.Tuning
Two non-obvious properties worth knowing before tuning:
- No positive
rrf_klets one leg’s rank-1 outrank two legs’ rank-2. If you want a single leg to dominate, usealpha, notrrf_k. Atalpha=0oralpha=1the dead leg’s retrieval is skipped entirely. rrf_kis only meaningful as deep as the legs actually retrieved. Atwindow_mult=3andtop_k=10each leg is 30 deep, and every value ofrrf_kabove ~28 makes the same decision — sweepingrrf_kalone at a shallow window will look like the knob does nothing. Sweep the two together.
Reading Back the Config
The recorded BM25 config —k1, b, and the analyzer_version the corpus was indexed with — is reported on the index handle, or null when the index has no full_text field:
Constraints and Gotchas
full_text+ explicitfilterable: trueis an error, andfull_text+pattern: trueis an error always. If you need a field both text-searchable and exact-match filterable, store it twice under two names.- Designations are fixed for the index’s lifetime. They currently cannot be changed after creation.
order_bycannot be combined withtext. Text results rank by score.- Only single string values are analyzed. A
full_textfield holding a number, boolean, object, or a list of strings is skipped for that item — join multi-part text into one string before upserting. - Not currently supported: phrase queries,
NOT/exclusion, nested boolean expressions, non-English analysis, and dual-indexing a field as both analyzed and exact-match. - The analyzer is versioned. The pipeline version, stopword-list hash, and stemmer identity are stamped into the index at creation (see
analyzer_versionabove). Opening an index with a build whose analyzer differs is a loud error, not silent recall loss. - Multi-tenant guidance: use index-per-tenant. Corpus statistics are per index, so co-mingling tenants behind a
tenant_idfilter gives every tenant blended term rarity — BM25 statistics are deliberately global and never scoped to a filtered subset.
Common Errors
Create-time schema violations are rejected by request validation (HTTP422); query-time and cross-parameter violations are enforced by the index core (HTTP 400). The SDKs surface both as errors/exceptions with the reason.
API Reference
For more information on full-text and hybrid search, refer to the API reference:REST API: Query
Hybrid query on
/v1/vectors/queryREST API: Query Metadata
Full-text query on
/v1/vectors/query_metadataPython SDK Reference
API reference for
query() and query_metadata() in PythonJS/TS SDK Reference
API reference for
query() and queryMetadata() in JavaScript/TypeScriptGo SDK Reference
API reference for
Query() and QueryMetadata() in GoCreate Index (REST API)
Designating
full_text fields at creation