Learn / RAG in 7 lessons / Retrieval quality and hybrid search
Retrieval quality and hybrid search
Why pure semantic search misses obvious matches, what hybrid search fixes, and how re-ranking tightens the final top-k.
Where pure vector search falls short
Semantic embeddings are trained to capture meaning, which is exactly why they’re good at matching a question phrased differently from the source text. It’s also exactly why they’re weak at:
- Exact identifiers - error codes, SKUs, ticket numbers, acronyms. A rare token like
ERR_504_TIMEOUTdoesn’t move an embedding much; the vector still mostly represents the surrounding prose. - Negation and specificity - “return policy” and “no return policy” can embed close together, because they’re topically related even though they mean opposite things.
- Very short, keyword-style queries - a two-word query has little semantic content for the model to work with.
A plain keyword method like BM25 (term-frequency ranking with document-length
normalization - the same family of algorithm this site’s own chatbot uses, in
src/lib/chat/bm25.ts) is the mirror image: excellent at exact-token matches, poor at
matching different phrasing for the same idea.
Hybrid search: use both, merge the results
Hybrid search runs a query through both a keyword index and a vector index, then merges the two ranked lists into one. The standard merge technique is Reciprocal Rank Fusion (RRF), which sidesteps the problem that BM25 scores and cosine-similarity scores aren’t on comparable scales:
rrf_score(doc) = sum over each ranker of 1 / (k + rank_in_that_ranker)
k is a small constant (commonly 60) that dampens the effect of rank 1 vs rank 2 while still
rewarding documents that rank highly in either list. A document that’s #1 in the keyword
list and unranked in the vector list still surfaces; so does one that’s #2 in both. No score
normalization required - only rank position matters.
Re-ranking: a second, more expensive pass
Retrieval (keyword or vector) has to be fast, because it searches the whole index. That speed comes from comparing lightweight representations. A re-ranker - usually a cross-encoder model that reads the query and one candidate chunk together - is slower per comparison but far more precise, because it can reason about the actual relationship between the two texts instead of comparing pre-computed vectors.
The pattern: retrieve a wider candidate set cheaply (say, top 20-50 from hybrid search), then re-rank only that shortlist down to the top 3-8 that actually go in the prompt. This keeps the expensive step bounded to a small, fixed number of comparisons per query instead of scaling with the size of the index.
What to actually do
- Start with vector search alone; it’s simpler and covers most natural-language questions.
- Add BM25 + RRF hybrid search once you notice exact-match queries (IDs, codes, acronyms) failing - a good sign is a support or search log full of short, specific queries.
- Add a re-ranker last, only if retrieval precision is still the bottleneck after hybrid search - it’s the most expensive lever, so pull it last, not first.
Key takeaways
- Pure vector search can miss exact matches - an error code, a product SKU, an acronym - because embeddings favor meaning over exact tokens.
- Hybrid search combines a keyword method (BM25) with vector search and merges the two ranked lists, catching both kinds of query.
- Reciprocal Rank Fusion (RRF) is a simple, effective way to merge two ranked lists without needing comparable raw scores.
- A re-ranker (a small cross-encoder model) can improve precision on the merged top-k before it goes into the prompt, at extra latency cost.
Quick check
3 questions - see how much stuck.