Text-search data is stored encrypted: the backing store never sees a readable term, which terms a document contains, or how many documents contain a given term.
How it works
- BM25 is enabled by designating fields, not by a flag. An index with at least one
full_textfield supports text search. An index with none is unaffected. - Text is sourced from metadata field values, not from
contents.contentsstays arbitrary binary and is never analyzed. - English only (currently): lowercase → split on non-alphanumeric → stopword filter → Porter2 stem.
- Designations are fixed for the index’s lifetime. They cannot currently be changed after creation.
- Fusion is by rank, not score. BM25 relevance and vector distance have no common unit, so hybrid queries fuse positions:
score(d) = Σ w_leg / (k + rank_leg(d)).alphasets the leg weights,rrf_kdamps how much a single leg’s top hit dominates, andwindow_multsets how deep each leg retrieves before fusing.
Designating fields
At index creation, mark fields asfull_text in metadata_schema — or use the text_fields shorthand:
full_text=True requires filterable=False, and is incompatible with pattern=True. In Python, {"full_text": True} alone implies filterable: False — only writing filterable: True yourself is a contradiction. If a field needs to be both searchable and exact-match filterable, store it twice under two names.Ingest
Nothing special is needed at upsert — text comes from the designated fields’ metadata values:full_text field that is missing or whose value is not a JSON string is skipped for that item (the value is never logged) and excluded from that field’s document count. An empty or all-stopword value is legal and simply contributes no terms.
Full-text query — no vector
score is present only when there is one — a filter-only query_metadata() call has nothing to score, so the key is absent rather than None.
Narrowing the searched fields, weighting, and AND
Composing with a filter
A filter combines with text as a pre-filter — it resolves first, then only its survivors are scored. Document frequency and per-field corpus stats stay corpus-global, so a document’s score never depends on the filter:order_by is not supported together with text — text results rank by score.
Hybrid query — BM25 + vector
Settext=... on query() to fuse BM25 with vector similarity:
score (fused relevance, larger = better) and never distance — a document matched by text alone has no vector distance to report.
A query vector is required even at alpha=0 — it fixes the batch shape. For text search with no vector, use query_metadata(text=...).
Hybrid queries narrow which filters are accepted.
$regex on a non-pattern field, and any predicate on an explicitly non-filterable field, work in a pure query() but are rejected once text is set, because the text leg can only be filtered on fields resolvable from the encrypted metadata index.Tuning
Two non-obvious properties worth knowing before tuning:
- No positive
rrf_klets one leg’s rank-1 outrank two legs’ rank-2. Solving1/(k+1) = 2/(k+2)givesk = 0. If you want a single leg to dominate, usealpha, notrrf_k. rrf_kis only meaningful as deep as the legs actually retrieved. Atwindow_mult = 3andtop_k = 10each leg is 30 deep, and every value ofkabove ~28 makes the same decision — so sweepingrrf_kalone at a shallow window will look like the knob does nothing. Sweep the two together.
Constraints and gotchas
full_text+filterable=True(explicit) is an error, andfull_text+pattern=Trueis an error always — a field can’t be both full-text and pattern-indexed. If you need a field both searchable and exact-match filterable, store it twice under two names.- Designations are fixed for the index’s lifetime. They cannot currently be changed after creation.
order_bycannot be combined withtext. Text results rank by score.- Not currently supported: phrase queries,
NOT/exclusion, nested boolean expressions, non-English analysis, and dual-indexing a field as both analyzed and exact-match. - The analyzer is versioned. The pipeline version, stopword-list hash, and stemmer identity are stamped into the index at creation. Opening an index with a build whose analyzer differs is a loud error, not silent recall loss — every existing document was indexed by the old analyzer.
- Multi-tenant guidance: use index-per-tenant. Corpus statistics are per index, so co-mingling tenants behind a
tenant_idfilter gives every tenant blended term rarity — BM25 statistics are deliberately global and never scoped to a filtered subset.