RAG Evaluation Metrics and Benchmarks
RAG evaluation metrics help teams tell whether a retrieval-augmented generation system finds the right evidence, uses it faithfully, and gives a useful answer. A polished response can still be wrong if retrieval missed a key passage, while a strong retrieval score can hide an answer that invents details. A useful benchmark separates these stages so engineers can locate failures before changing prompts, indexes, or models.
What Is RAG Evaluation?
Retrieval-augmented generation (RAG) answers questions by retrieving passages from a chosen collection and supplying them to a language model. Evaluation is the process of measuring whether that pipeline behaves as intended across representative questions, rather than relying on a few demonstrations or subjective impressions.
The Ragas project provides tools for evaluating RAG systems, and its metric documentation describes measures for both retrieved context and generated answers. These metrics are useful, but they are not a universal pass/fail standard: teams still need a test set that reflects their users, sources, and consequences of error. For an overview of the pipeline itself, see retrieval-augmented generation best practices.
RAG evaluation usually combines two kinds of evidence. Retrieval metrics compare the passages returned by a search system against passages judged relevant to a query. Answer metrics check properties such as whether the response is supported by those passages or satisfies a reference answer. Human review and operational measurements, including latency and cost, complete the picture.
The Problem RAG Evaluation Solves
Without a stable evaluation set, teams often judge a RAG change by trying a few familiar questions. That can miss regressions that affect less common queries, exact identifiers, recently changed documents, or requests that should be refused. A prompt adjustment may make answers sound clearer while causing the retriever to miss the source that supports them.
The pipeline also has several distinct failure points. Relevant material may not have been indexed, may rank below the context limit, or may be retrieved but omitted during context assembly. The model may then misread evidence, combine unrelated passages, or answer when the evidence is insufficient. A single satisfaction score cannot identify which stage failed.
A benchmark creates a repeatable comparison between a baseline and a change. It can show whether a new embedding model improves recall but increases latency, whether a reranker helps critical queries, or whether a prompt reduces unsupported claims. It does not prove that a system is safe or correct for every future question; it makes observed behavior measurable and easier to investigate.
How RAG Evaluation Works
An evaluation example should represent one meaningful user task. At minimum, record the question and the relevant source passages. For answer evaluation, also record the expected answer or required facts, and note whether the system should abstain. Include query types, source versions, and access conditions when they matter to the product.
Run each example through the system and retain intermediate results: retrieved passage IDs and ranks, the context sent to the model, the generated answer, and any citations. Then score retrieval and generation separately. Compare candidate changes on the same examples, inspect results by query category, and review failures instead of treating the aggregate average as the whole result.
The NIST AI Risk Management Framework is broader than RAG and does not prescribe these scores, but it emphasizes managing and measuring AI risks. That distinction matters: benchmark metrics are evidence for decisions, not a replacement for risk analysis, human oversight, or product-specific acceptance criteria.
Key Metrics and Components
Different metrics answer different questions. The table summarizes commonly used measures; the exact definition and judging method should be documented for each benchmark.
| Metric | Stage | What it measures | Main limitation |
|---|---|---|---|
| Precision@k | Retrieval | Share of the top k results that are relevant | Does not penalize relevant passages that were not retrieved |
| Recall@k | Retrieval | Share of known relevant passages found in the top k | Requires a useful set of relevance labels |
| Mean reciprocal rank (MRR) | Retrieval | How high the first relevant result appears, averaged over queries | Ignores other relevant results after the first |
| nDCG@k | Retrieval | Ranking quality when relevance has multiple grades | Depends on consistent graded labels and a chosen cutoff |
| Context precision and recall | Retrieval/context | Whether relevant context is prioritized and whether it contains the needed information | Often relies on an LLM judge or reference annotations |
| Faithfulness or groundedness | Answer | Whether answer claims are supported by the supplied context | A response can faithfully repeat a false or outdated source |
| Answer correctness and relevance | Answer | Whether the response matches expected facts and addresses the question | Reference answers can be incomplete; judge scores need calibration |
| Latency, cost, and task success | Whole system | Whether the system meets operational and user-task requirements | These are not captured by a retrieval or answer-quality score |
Precision and recall depend on a clear definition of relevance. For a question about a specific configuration option, the exact documentation page may count as relevant; for a policy question, several passages may each provide partial evidence. If annotations are incomplete, recall can unfairly penalize a system for returning a valid passage that annotators did not label.
Ranking measures add position information. MRR rewards placing the first relevant result near the top. nDCG can represent graded usefulness, such as a directly answering passage being more valuable than a passage that only provides background. Both need an explicit cutoff because a RAG system typically sends only some retrieved passages to the model.
Answer-level evaluation needs equally careful interpretation. Faithfulness asks whether claims are supported by the supplied context, not whether that context is true. Correctness can be checked against human-written facts or reference answers, while answer relevance asks whether the response addresses the request. LLM judges can help scale these checks, but their rubric, model, and known disagreements should be tracked; a judge score is not ground truth.
Real-World Use Cases
An internal support assistant can be tested for whether it retrieves the current approved procedure and avoids answering when the procedure is missing. A developer documentation assistant can be evaluated separately on conceptual questions and exact symbols, version numbers, or error messages. Those slices often expose why hybrid search for RAG may help some queries without improving every query.
For regulated or high-impact workflows, include questions with conflicting, outdated, or insufficient evidence and define when a person must review the result. A benchmark should also test permissions: a relevant passage that the user cannot access is not a successful retrieval. Production traces can reveal new query patterns, but sensitive data should be handled according to the system’s privacy and retention rules before it is used for evaluation.
Getting Started: Build a Small Retrieval Benchmark
Start with a modest set of real or carefully written questions, then label the relevant passage IDs for each one. Add expected answer facts or an abstention label for tasks where the output must be checked. Have a domain reviewer inspect a sample of labels, keep a held-out set for final comparisons, and version the documents and labels so a score can be reproduced.
The following Python example calculates Precision@k, Recall@k, reciprocal rank, and binary nDCG for one query. It uses only the Python standard library, so no package installation is required. The relevant IDs must match the same unit returned by retrieval, such as chunks or documents.
from math import log2
def evaluate_retrieval(retrieved_ids, relevant_ids, k=3):
if k < 1:
raise ValueError("k must be at least 1")
if not relevant_ids:
raise ValueError("relevant_ids must contain at least one known relevant item")
top_k = retrieved_ids[:k]
hits = [item_id in relevant_ids for item_id in top_k]
hit_count = sum(hits)
first_hit_rank = next(
(rank for rank, is_hit in enumerate(hits, start=1) if is_hit),
None,
)
reciprocal_rank = 1 / first_hit_rank if first_hit_rank else 0.0
dcg = sum(
int(is_hit) / log2(rank + 1)
for rank, is_hit in enumerate(hits, start=1)
)
ideal_hit_count = min(len(relevant_ids), k)
ideal_dcg = sum(
1 / log2(rank + 1)
for rank in range(1, ideal_hit_count + 1)
)
return {
f"precision@{k}": hit_count / k,
f"recall@{k}": hit_count / len(relevant_ids),
"reciprocal_rank": reciprocal_rank,
f"ndcg@{k}": dcg / ideal_dcg,
}
retrieved = ["doc-17", "doc-4", "doc-8"]
relevant = {"doc-4", "doc-9"}
for metric, score in evaluate_retrieval(retrieved, relevant).items():
print(f"{metric}: {score:.3f}")
Save the script as evaluate_retrieval.py, then run python --version and python evaluate_retrieval.py in a terminal. This example is a retrieval check, not a complete RAG benchmark: it does not score answer support, correctness, latency, or access control. For a fuller evaluation, add those measures to the same versioned test cases, inspect the lowest-scoring examples, and rerun the unchanged baseline after each system change.
Common Misconceptions
- “One score proves the system is good.” Averages can hide weak query categories, permission failures, or rare but costly errors. Keep stage-level metrics and inspect representative failures.
- “Faithfulness means correctness.” A model can accurately quote a stale or incorrect source. Source quality and freshness need their own checks.
- “An LLM judge makes evaluation objective.” Judges can be inconsistent or favor particular answer styles. Define a rubric, compare a sample with human ratings, and record the judge version.
- “A benchmark stays representative forever.” User questions and document collections change. Review coverage periodically, add new failure cases, and preserve a held-out set so continuous updates do not turn the benchmark into a memorized target.
Related Articles
- Retrieval-Augmented Generation Best Practices covers the broader design and quality controls of a RAG pipeline.
- Hybrid Search for RAG: Combining Vector and Sparse Retrieval explains how to compare retrieval strategies.
- Vector Databases for RAG: How Retrieval Works describes embeddings, indexes, and retrieval behavior.
- OpenTelemetry for LLM Production Observability covers traces and operational signals for deployed model systems.
Changelog
- Initial publication.
Last updated: September 30

