The Context Window

Hybrid Search Combining Dense and Sparse Retrieval

Two retrieval methods catch what the other misses, boosting accuracy over either alone.

Reporter · · 8 min read
Cover illustration for “Hybrid Search Combining Dense and Sparse Retrieval”
Context Delivery · September 26, 2026 · 8 min read · 1,868 words

Hybrid search means running a query through two different retrieval methods at once, sparse keyword matching and dense vector search, then merging the two ranked lists into one. It works because the two methods' failures point in opposite directions. Where one goes blind, the other tends to see clearly, and that predictable asymmetry is what makes fusion worth the extra engineering.

Limitations of pure keyword search and pure vector search on real queries

Any real corpus, whether it's a support knowledge base or a product catalog, gets hit with two genuinely different kinds of queries. Some are semantic: a user types "how to cancel" when the document says "subscription termination process," and the words barely overlap even though the intent is identical. Others are exact-match by nature, such as a SKU number, an error code, an API method name, a regulatory clause citation, or a person's name. These are queries that succeed or fail based on completely different mechanics. They're queries that succeed or fail based on completely different mechanics.

Dense retrieval handles the first kind well and the second kind poorly, and it fails silently. Embeddings encode meaning, not exact strings, so a rare token or an out-of-vocabulary identifier gets pulled toward some generic region of the embedding space instead of standing apart from it. Nothing crashes. There's no error thrown, no latency spike to flag the miss. The retrieved chunks are semantically adjacent to the query but lexically wrong, and the language model downstream writes a fluent, confident answer anyway, built on the wrong material. On mixed-query benchmarks, pure dense retrieval is around 0.72 NDCG@10 on mixed benchmarks.

Sparse retrieval fails the opposite way. BM25 matches tokens, full stop, so "bicycle repair" retrieves nothing useful when the source document talks about "fixing a bike." No amount of relevance exists in the term frequency math to bridge that gap, because the vocabulary itself doesn't overlap. On the same mixed-query corpora, pure BM25 is around 0.58 NDCG@10, a meaningfully lower ceiling than dense retrieval's, driven entirely by the semantic queries it can't touch.

The asymmetry here is structural. It's structural. Dense retrieval loses on exact terms; sparse retrieval loses on paraphrase and vocabulary mismatch. Combining the two is corrective, patching each method's blind spot with the other's strength. It's corrective, patching each method's blind spot with the other's strength.

Where BM25's precision comes from

BM25 runs on an inverted index, a data structure that, for every term appearing anywhere in the corpus, stores the list of documents containing that term along with how often it occurs. At query time, the engine looks up each query term in that index, scores every candidate document against it, and ranks the results. No neural network runs. No embeddings get computed. It's counting and arithmetic.

Two scoring mechanisms do the real work. Term frequency saturation means that additional occurrences of a word contribute less and less to the score as they pile up, which stops a document from gaming its way to the top simply by repeating a keyword fifty times. Length normalization penalizes longer documents, since a longer document naturally contains more term matches by sheer volume and would otherwise win on that basis alone rather than on relevance.

Both mechanisms are tunable. The parameter k1 controls term frequency saturation and b controls length normalization, and the defaults, k1=1.2 and b=0.75, work reasonably well across general text. But they're defaults, not laws. Short structured documents, like entries in a product catalog, may benefit from adjusting k1 downward, since repeated terms in a short record are less informative than they'd be in a long essay. Long-form documents where genuine term repetition signals relevance benefit from pushing k1 toward the higher end of its range instead.

What BM25 does well, it does very well: exact-match queries, terms an embedding model has simply never encountered, and high-throughput retrieval with no neural inference required at query time. Query latency is low, driven by hash table lookups and integer math, entirely on CPU. This runs entirely on CPU, with no GPU involved.

Dense retrieval and semantic matching in practice

Dense retrieval starts with a transformer encoder that maps both the query and every document chunk into high-dimensional vectors, computed once at index time for documents and at query time for the incoming query. Retrieval then becomes a similarity search, cosine similarity or dot product, over that vector space. Approximate Nearest Neighbor indexes, HNSW and FAISS among the most widely used, make that search fast enough to run in production instead of scanning every vector by brute force. In a hybrid pipeline, the embedding inference step itself, not the search operation, tends to dominate total query cost, often accounting for a substantial share of the latency.

"Automobile" retrieves documents about "cars" and "vehicles" because their vectors land close together in embedding space, regardless of the fact that the strings share no characters. On open-domain question answering, DPR reaches 75.4% Top-1 accuracy on Natural Questions, against BM25's 54.0% on the same benchmark, a wide enough gap to make the semantic advantage concrete rather than theoretical.

That advantage comes with real operational cost. GPU-backed indexing runs meaningfully more expensive than building an inverted index, and swapping in a new embedding model isn't a config change, it means reindexing the entire corpus from scratch offline.

BGE-M3, from BAAI, is a useful reference point for what a modern dense retrieval backbone actually looks like. It's a multilingual, multi-functionality bi-encoder that supports dense, sparse, and multi-vector retrieval at once, built on XLM-RoBERTa-large (24 layers, 1024 hidden dimensions, 16 attention heads) and enhanced with RetroMAE pretraining, landing around a substantial number of parameters. It handles input sequences up to 8,192 tokens, which matters for retrieval over long documents rather than short snippets, and covers over 100 languages, which is a large part of why it appears so often as the embedding backbone in current RAG pipelines. Inference on a 512-token query runs 5 to 10 milliseconds on an A100 GPU. In a financial information retrieval benchmark (arXiv 2511.00855), BM25 scored NDCG@5 of 0.1768 while BGE-M3's sparse mode reached 0.5722, a gap that shows learned, dense-enhanced sparse representations can beat classical lexical scoring by a wide margin even while staying inside the "sparse" retrieval family.

How fusion algorithms merge two ranked lists into one

Fusion exists to solve one specific problem: a cosine similarity of 0.85 from vector search and a BM25 score of 12.4 sit on two completely different, mathematically unrelated scales. You cannot average them, weight them, or compare them directly without first deciding what to do about the fact that one is bounded between -1 and 1 and the other is an unbounded log-scale term-frequency score.

Reciprocal Rank Fusion, RRF, is the production default for a reason. Introduced by Cormack and colleagues, its formula scores each document as the sum, across both ranked lists, of 1 divided by (k plus that document's rank in each list). Because it works entirely on rank position rather than raw score, it avoids the scale mismatch between the two ranked lists completely, no calibration step needed. The constant k=60 is the commonly used default, and RRF ships as the native retriever option in Elasticsearch and as the out-of-the-box default in Microsoft Azure AI Search. Documents that rank highly in both the sparse and dense lists accumulate the highest combined RRF scores, so the method naturally rewards agreement between the two retrieval signals rather than just picking a winner. Later research has generally confirmed that RRF is difficult to beat without collection-specific tuning, particularly when there isn't enough training signal available to justify a more elaborate approach.

Alpha-weighted score fusion is the tunable alternative, and it works differently: score(d) = α × dense_score + (1 − α) × sparse_score, with α set by hand. Weaviate defaults to α=0.5, an even split between dense and BM25 signal, and LlamaIndex exposes alpha as a first-class tuning parameter rather than burying it. Left untuned, BM25's unbounded scores tend to dominate the blend simply because they're numerically larger than cosine similarity's bounded range, which can quietly skew results toward keyword matches even when that's not the intent. Pinecone's documented approach is to just test it: run several alpha values, say 0.25, 0.5, and 0.75, against a real query set and pick whichever maximizes retrieval quality on that specific data. There's no universal optimum sitting in a table somewhere.

And no fusion method wins everywhere. That's an empirical result. Domain characteristics, acronym density, how stable the vocabulary is between queries and documents, the mix of exact-match versus conceptual queries, all shift which fusion method comes out ahead. Choosing RRF by default is reasonable. Assuming it's always correct is not.

The extent of hybrid retrieval's benefit

E-commerce retrieval is close to the ideal use case, because product catalog queries mix exact SKU lookups with semantic style and category searches in the same query stream, stressing both failure modes simultaneously. A tuned hybrid setup reaches 0.7497 NDCG, a 7.4% lift over BM25 alone (0.6983) and over pure vector search alone (0.6953). Neither single method gets close on its own.

Financial documents with mixed text and tables show an even larger effect, a two-stage pattern. Hybrid RRF alone reaches Recall@5 of 0.695, against BM25 alone at 0.644 and dense alone at 0.587, already a clear win for the hybrid approach. But adding a reranking stage on top, Cohere Rerank v4.0 Pro layered after hybrid retrieval, pushes Recall@5 to 0.816: a 17.4% gain over hybrid RRF alone, 26.7% over BM25, and 39.0% over dense retrieval alone. MRR@3 moves from 0.433 with hybrid RRF to 0.605 with reranking added, a 39.7% relative jump. That pipeline used text-embedding-3-large for the embedding step and Cohere Rerank v4.0 Pro for reranking, both served through Azure AI Foundry, and the scale of the reranking lift suggests fusion alone, however well-tuned, still leaves value on the table that a dedicated reranking pass can capture.

Scientific literature tells a different story, and it's an important corrective. In a silicon detector R&D benchmark (arXiv 2606.24725), hybrid retrieval reaches Hit@5 of 0.917 on the core benchmark and 0.951 on the extension benchmark, strong numbers on their face. But the gap between hybrid and BM25 alone in this domain is notably narrow. Scientific corpora tend to be acronym-rich and terminologically stable, so the vocabulary in a researcher's query and the vocabulary in the paper largely overlap already, BM25 has little semantic gap left to fail on, and hybrid's marginal contribution shrinks accordingly.

RAG-Fusion (Raudaschl) offers one more data point, combining multiple query reformulations with RRF fusion across both BM25 and vector search for each reformulated query. The Hybrid+Diverse configuration produced a 19% gain in NDCG@10 and an 18% gain in MRR over baseline, suggesting that the gains reflect a real improvement rather than noise.

Read across all four domains, the pattern holds: hybrid retrieval's advantage over either pure method scales with how much the query vocabulary and the document vocabulary diverge. Where that divergence is high, e-commerce, financial documents, mixed exact-match and conceptual queries, hybrid retrieval earns its added engineering complexity. Where vocabulary is already shared and stable, as in specialist scientific literature, BM25 alone holds up competitively, and the case for hybrid gets thinner.

Diagram: Hybrid Retrieval Closes the Gap Both Methods Leave Open. Visualizes: Show the performance contrast between pure BM25, pure dense retrieval, and hybrid retrieval across two retrieval metrics, using actual benchmark numbers from the article.

Sources

  1. Hybrid Search for RAG: Combining BM25 and Dense Vector Search (2026 Guide)
  2. Hybrid RAG: Dense and Sparse Retrieval for Better AI Answers
  3. arxiv.org
Filed underContext Delivery

More in Context Delivery