Hybrid Search for RAG: Combining Vector and Sparse Retrieval

Updated on
10 min read

Hybrid retrieval combines dense vector search with sparse, keyword-oriented search so a system can find both conceptually related passages and documents containing exact terms. It is useful for teams building retrieval-augmented generation (RAG), enterprise search, or documentation search where a query might mention an error code in one request and describe a concept in another. This guide explains the two retrieval paths, how to combine their rankings, and how to evaluate whether the added complexity improves results.

What Is Hybrid Retrieval?

Hybrid retrieval runs more than one search method for a query and combines their results into a final ranked list. A common design pairs dense retrieval, which compares embeddings to find semantic similarity, with sparse retrieval, which represents text using weighted terms and matches those terms through an inverted index.

The OpenSearch project and its hybrid search documentation describe a hybrid query that runs multiple query clauses and combines their scores through a search pipeline. Other search systems expose the same general pattern: Weaviate’s hybrid search, for example, fuses vector and BM25F results and lets applications adjust their relative influence. The product APIs differ, but the system-level idea is portable.

Hybrid search is not a new kind of embedding or a guarantee of better answers. It is a retrieval architecture: independent methods produce candidate documents, a fusion step orders them together, and an application selects evidence for its next stage.

The Problem Hybrid Retrieval Solves

Each search method has a different failure mode. Dense vectors can connect paraphrases even when their words differ, but may rank an approximate semantic match above an exact identifier. A query for ERR_CONN_RESET or a model name may need literal token matching. Lexical search such as BM25 is good at those terms, but a query phrased differently from the source may fail to retrieve a relevant passage.

In RAG, a missed passage can affect the generated answer: the model can only use evidence that reaches its context. A lexical index alone can overlook paraphrases; a vector index alone can overlook rare names, numbers, or exact phrases. Running both paths gives a system a broader candidate pool, particularly for technical corpora that mix prose, product names, code, and numeric values.

This coverage has a cost. Two indexes consume more storage and must be updated consistently. Each query does more work, and scores from different retrievers are not necessarily on the same scale. Hybrid retrieval is therefore a trade-off to measure, not a default upgrade.

How Hybrid Search Works

A typical hybrid query follows these steps:

  1. Prepare the corpus. Extract and chunk documents, preserve identifiers and access metadata, index text for lexical search, and create embeddings for dense search.
  2. Run retrieval paths. Search the same authorized document set with a sparse query and an embedding query. Each path returns a ranked candidate list, often with its own relevance score.
  3. Fuse candidates. Combine scores or rankings, remove duplicate documents, and produce a single ordered list.
  4. Optionally rerank. A more expensive model can score a limited candidate set using both the query and each passage.
  5. Assemble context. Apply access checks, choose a small set of useful passages, retain source identifiers, and send them to the language model.

The key is that the two result lists are not directly interchangeable. A BM25 score is based on lexical statistics; a vector similarity score depends on an embedding model and distance metric. Adding raw scores without accounting for their scales can let one retrieval path dominate.

Two common fusion approaches address this:

  • Score normalization and weighted combination first transform scores to comparable ranges, then combine them. A weighted mean can express a preference for exact keyword matches or semantic similarity. The OpenSearch normalization processor supports normalization methods and combination weights for hybrid queries.
  • Reciprocal rank fusion (RRF) combines result positions rather than raw scores. A document’s contribution from one list is commonly expressed as 1 / (k + rank), where k is a smoothing constant. Contributions from each list are added, so a document ranked highly by multiple methods tends to rise in the merged list. Elasticsearch’s RRF reference describes the approach and its parameters; the original RRF paper evaluated rank fusion across search tasks.

RRF avoids calibrating the score scales, while a normalized weighted sum allows explicit score weighting. Neither fusion method can recover a relevant document that both retrievers failed to return. Candidate depth, filters, and corpus quality still matter.

Key Retrieval Components and Choices

Component What it searches or changes Strength Main trade-off
Dense vector retrieval Embeddings ranked by a similarity or distance measure Finds related meaning and paraphrases May miss exact identifiers; depends on embedding quality
Lexical sparse retrieval Token matches through an inverted index, often scored with BM25 Matches rare words, names, numbers, and phrases Different wording can hide relevant content
Learned sparse retrieval Model-produced weighted token representations in a sparse index Can add semantic expansion while retaining term-level matches Adds model, index, and serving complexity
Score fusion Normalized scores combined with configured weights Makes relative contribution adjustable Requires calibration; score distributions can shift
Rank fusion (RRF) Positions in separate result lists Combines rankings without comparing raw score scales Ignores score gaps and still needs tuning of candidate depth

Sparse does not always mean traditional keyword search. BM25 operates on lexical terms and corpus statistics. Learned sparse models can assign weights to a larger vocabulary, including terms related to a query or document, while still producing sparse representations. Dense and sparse retrieval are representation choices; hybrid describes combining outputs from multiple paths.

Candidate depth and filtering affect what can be fused. If each path returns only a few results, useful passages may never reach the fusion stage. If one path returns many weak candidates, latency and reranking cost may rise. Apply tenant and permission filters in each retrieval path before results can enter the merged list; filtering only after retrieval can leak restricted passages into logs, reranking services, or downstream processing.

Reranking is a separate optional stage, not a synonym for fusion. A cross-encoder can compare the query and each candidate together, often improving the top ordering at the cost of more computation. Keep its candidate set bounded and evaluate it separately from the first-stage retrieval.

Real-World Use Cases

  • Technical support: Sparse search can recover exact error codes and product versions, while dense search finds troubleshooting documents that describe the same failure in different words.
  • Internal knowledge assistants: Hybrid retrieval can handle both a question such as “how do I rotate a key?” and a query containing an exact policy name. Authorization filters must follow the user’s access rights in both indexes.
  • Code and API documentation: Exact symbols and endpoint names are important, but developers may describe what a function should do rather than know its name.
  • Product or catalog search: A person may use an exact model number, a descriptive phrase, or both. Separate retrievers can surface candidates from each type of intent.

These systems should preserve provenance through fusion. A useful result needs a stable document and chunk identifier so the application can cite it, check freshness, and inspect why a passage was returned.

Practical Guide: Configure a Small OpenSearch Hybrid Query

The following local example uses OpenSearch’s text and k-NN query clauses with a normalization pipeline. It is for development only: the security plugin is disabled, so do not expose this container to an untrusted network. For a local Docker environment, start a single node:

docker run -d --name opensearch \
  -p 9200:9200 -p 9600:9600 \
  -e "discovery.type=single-node" \
  -e "DISABLE_SECURITY_PLUGIN=true" \
  -e "OPENSEARCH_JAVA_OPTS=-Xms1g -Xmx1g" \
  opensearchproject/opensearch:2.19.0

After the container starts, check that the node responds:

curl http://localhost:9200

Create an index with a text field and a three-dimensional vector. The small vector is only for demonstrating the request shape; production vectors must come from one embedding model and use its actual dimension:

PUT /knowledge
{
  "settings": {
    "index": {
      "knn": true
    }
  },
  "mappings": {
    "properties": {
      "text": { "type": "text" },
      "embedding": {
        "type": "knn_vector",
        "dimension": 3
      }
    }
  }
}

Add one document with a matching vector:

PUT /knowledge/_doc/1
{
  "text": "Rotate an API key from the account security settings.",
  "embedding": [0.1, 0.2, 0.3]
}

Configure score normalization and equal weighting for the two query clauses. The weights follow the order of the hybrid.queries array:

PUT /_search/pipeline/rag-hybrid
{
  "phase_results_processors": [
    {
      "normalization-processor": {
        "normalization": { "technique": "min_max" },
        "combination": {
          "technique": "arithmetic_mean",
          "parameters": { "weights": [0.5, 0.5] }
        }
      }
    }
  ]
}

Run both searches through the pipeline. Replace the example vector with the query embedding generated by the same model used for the indexed documents:

GET /knowledge/_search?search_pipeline=rag-hybrid
{
  "size": 5,
  "query": {
    "hybrid": {
      "queries": [
        { "match": { "text": { "query": "change the API key" } } },
        { "knn": { "embedding": { "vector": [0.1, 0.2, 0.3], "k": 5 } } }
      ]
    }
  }
}

For a production check, create a test set of representative queries with known relevant passages. Compare dense-only, sparse-only, and hybrid results using retrieval measures such as Recall@k, mean reciprocal rank, or nDCG@k. Also record query latency, index size, and any embedding or reranking costs. Inspect failures by query type: if hybrid helps exact identifiers but hurts conceptual questions, adjust fusion weights or investigate each retriever rather than assuming one global setting is best.

Those measures cover search quality, but not whether generated answers are supported by retrieved evidence. Retrieval-augmented generation best practices covers evaluation across retrieval and answer generation.

Common Misconceptions

  • “Hybrid search means using two vector databases.” It means combining retrieval signals; one search engine can host both text and vector indexes, or separate systems can be orchestrated by an application.
  • “Hybrid always beats dense or keyword search alone.” It can improve coverage, but may add latency, storage, and irrelevant candidates. A representative evaluation set should decide whether the result is worthwhile.
  • “Fusion fixes a weak retriever.” Fusion only combines returned candidates. Improve document extraction, chunking, query handling, or index configuration when relevant content is absent from both candidate lists.
  • “A reranker makes retrieval hybrid.” Reranking changes candidate order; it does not necessarily combine independently retrieved sparse and dense result sets.

Changelog

  • Initial publication.

Last updated: September 29, 2026

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.