Blog/RAG & retrieval

Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed

How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.

··12 min read

  • RAG evaluation
  • Retrieval metrics
  • LLM-as-judge
  • Golden set
  • Ragas
Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.

Key takeaways

  • A RAG system has two halves that fail differently, so score them separately: retrieval with recall@k, MRR and nDCG, generation with faithfulness and answer relevance.
  • Recall@k is the ceiling. If the needed chunk is not in the top k, no prompt or model can produce a grounded answer, so check retrieval first when an answer is wrong.
  • A golden set of real user queries beats synthetic questions. Start small, label what a correct answer needs, and grow it from every production failure.
  • An LLM judge is a measuring instrument, not a ground truth. Calibrate it against human labels on a sample, report chance-corrected agreement and re-check it whenever the judge model changes.
  • Run cheap retrieval checks on every pull request and the LLM-judged suite on a schedule or when the index, prompt or model changes, and gate on regression against a baseline.

Most RAG systems are shipped on the strength of a demo. Someone asks five questions, the answers look right, and the team moves on. Then the corpus grows, the chunking changes, the model is upgraded, and nobody can say whether the system got better or worse. The honest answer to "is it good?" is a number you can re-compute, and for RAG that takes more than one number.

The reason is structural. A RAG system is two systems in a row: a retriever that decides what the model sees, and a generator that decides what to say about it. Each fails in its own way, needs its own metrics and gets fixed by different work. If you only score the final answer, a wrong answer tells you that something broke, but not where.

This article is how I would set up RAG evaluation for a product team: the metrics and what each one actually measures, a golden set built from real queries, an LLM judge you have checked against humans, the tools worth looking at, evals that run in CI, and a short procedure for diagnosing whether retrieval or generation failed. It builds on my posts about evals for LLM product features and the RAG pipeline itself.

Two halves, seven numbers

The RAGAS paper frames RAG evaluation along three dimensions: the retrieval system's ability to find relevant and focused passages, the generator's ability to use those passages faithfully, and the quality of the output. In practice I split the numbers by the question they answer. Retrieval metrics ask did the right material reach the model, and in a good order? Generation metrics ask given what it saw, did the model behave?

The table lists the metrics I actually use. Names differ between tools, so I describe what each one measures rather than what a library calls it.

MetricLayerThe question it answersNeeds labels?
recall@kRetrievalIs the relevant material anywhere in the top k results?Relevant chunks per query
MRRRetrievalHow high is the first relevant result? The mean of 1 divided by its rank.Relevant chunks per query
nDCGRetrievalIs the whole ranking in a good order, with graded relevance and a discount for lower positions?Graded relevance
Context precisionRetrieval (ranking)Are relevant chunks ranked above irrelevant ones in what the model received?Optional reference answer in Ragas
Context recallRetrievalIs everything the reference answer says supported by the retrieved context?Reference answer
FaithfulnessGenerationIs each claim in the answer supported by the retrieved context?No
Answer relevanceGenerationDoes the answer address the question that was asked?No

Note that context precision sits with retrieval, even though it is often listed beside the generation metrics. It is computed over the retrieved chunks, so it tells you about the ranker, not about the model's writing. Keeping the layer column honest is what makes the diagnosis later in this article possible.

Retrieval metrics: recall@k, MRR and nDCG

Retrieval metrics come from search evaluation and need one thing the generation metrics do not: a notion of which chunks are relevant for each query.

  • Recall@k asks whether the relevant material is in the top k. It is the most important retrieval number, because it is a ceiling. If the chunk that contains the answer is not among the k chunks you pass to the model, the model cannot ground its answer in it, whatever the prompt says. Choose k to match what you actually send to the generator.
  • MRR (mean reciprocal rank) averages 1 divided by the rank of the first relevant result: 1 for first place, 0.5 for second. It rewards putting a good chunk at the top. Its documented limitation is that only the first relevant result counts and any further ones are ignored, so it suits questions with one right passage and misleads on questions that need several.
  • nDCG sums graded relevance scores with a logarithmic discount for lower positions, then divides by the ideal ordering so the result lies between 0 and 1. It is the metric to use when relevance is not binary, for example a chunk that fully answers versus one that only mentions the topic, and when ordering matters because the context window or a reranker's cut-off truncates the list.

I read them together. High recall@k with a poor MRR or nDCG means the right chunk is retrieved but buried, which a reranker usually fixes (see the pipeline post for chunking, hybrid search and reranking). Low recall@k means the right chunk is never found, which is a problem with chunking, embeddings, query rewriting or the index, and no reranker can repair a candidate list that does not contain the answer.

These metrics need relevance labels, which is the expensive part. The Ragas documentation describes an ID-based and a non-LLM variant of context recall that compare retrieved chunk IDs or text to reference contexts without an LLM call. Those are cheap, deterministic and perfect for CI, if your golden set stores which chunks are relevant.

Generation metrics: faithfulness and answer relevance

Generation metrics judge the answer given what the model was shown. Faithfulness is the central one. In the Ragas definition it is the number of claims in the response that the retrieved context supports, divided by the total number of claims: the response is decomposed into statements, each statement is checked against the context, and the ratio is the score, from 0 to 1. It needs no reference answer, which is why it is so useful on production traffic.

Answer relevance asks whether the answer addresses the question at all. An answer can be perfectly faithful to the context and still dodge the question, for example by summarising a document when the user asked for a number. TruLens calls the equivalent triad context relevance, groundedness and answer relevance; Ragas calls it response relevancy; DeepEval calls it answer relevancy. Same idea, different labels.

Two properties are easy to miss. First, faithfulness measures agreement with the retrieved context, not with the truth. If retrieval returned a stale policy document, a perfectly faithful answer is wrong. Second, a high score can be earned by saying little: a model that refuses or hedges makes few claims, and few claims can all be supported. Always read faithfulness next to answer relevance and a correctness check against the golden answer.

Build the golden set from real queries

Everything above depends on a golden set: queries with the expectations you score against. Synthetic questions generated from your documents are a useful bootstrap, but they share the vocabulary of the documents, which real users do not. Real queries have typos, abbreviations, vague phrasing, product codes and questions the corpus cannot answer. I would build the set from them.

  1. Sample real queries. Take them from search logs, support tickets or chat transcripts. Remove personal data first, and mind your legal basis: see my notes on GDPR and LLM APIs.
  2. Stratify, do not just sample. Include the common questions, the long tail, multi-document questions, questions with exact identifiers, and questions that should get "I don't know" because the corpus has no answer.
  3. Label what a correct answer needs. For each query, record the relevant chunks or documents (for recall@k, MRR and nDCG), a short reference answer, and where it matters the facts the answer must include.
  4. Have a domain expert label it, and measure whether two people agree. If two humans disagree about what is relevant, no metric can resolve that for you.
  5. Version it and keep a held-out part. Treat the set like code. Keep a slice you never tune against, so improvements are not just overfitting to the examples you stared at.
  6. Feed it from production. Every thumbs-down, escalation or wrong answer found in review becomes a new case, with its label. This is how the set stays representative as the product changes.

How big? I would start with a few dozen well-labelled queries and grow from there; a small clean set that you trust beats a large one nobody checked. The point is not statistical power at the start, it is to have something that makes a regression visible the day it happens. As the set grows, report results per stratum, because an average hides a collapse in one query type.

Diagnose: did retrieval or generation fail?

When a wrong answer is reported, resist the urge to edit the prompt. Walk the chain from the bottom, using the stored retrieval results for that query, and stop at the first question you answer with "no".

Diagnosing a bad RAG answerA decision chain of four questions. If the fact is not in the corpus, it is a content gap. If it is in the corpus but not in the top k chunks, retrieval failed. If it is in the context but the answer is not supported, generation is unfaithful. If it is supported but does not address the question, the answer is off target. If all checks pass, re-check the golden label or the judge.Start at the top, stop at the first noIs the fact in the corpus?yesnoContent gapFix the sources, not the modelIs it in the top k chunks?yesnoRetrieval failurerecall@k, MRR, nDCG are lowIs every claim supported?yesnoGeneration: unfaithfulfaithfulness is lowDoes it answer the question?yesnoGeneration: off targetanswer relevance is lowAll pass: re-check label or judge
The order matters: each question assumes the ones above it have passed. Most of the surprises I see are in the first two.

The branches map to different work. A content gap is an editorial or ingestion problem: the document is missing, outdated or was never parsed properly. A retrieval failure is fixed in chunking, embeddings, hybrid search, query rewriting or reranking. An unfaithful answer is a generation problem: tighten the instruction to answer only from the context, reduce irrelevant context that invites speculation, or change the model. An off-target answer is usually the prompt, or a retrieval problem in disguise where the context is on topic but not on the question.

The last box matters too. If every metric passes and the answer is still wrong, suspect the golden label, the reference answer or the judge itself before the system. For agentic setups, where the model decides what to retrieve and may retrieve several times, the same logic applies per retrieval step; see RAG in 2026: hybrid, agentic and long context.

LLM-as-judge: calibrate it against humans

Faithfulness, answer relevance and most context metrics are scored by an LLM. Ragas notes that its LLM-based metrics may use one or more LLM calls to produce a score. That makes the judge part of your measurement setup, and an instrument needs calibration.

The evidence is encouraging but not unconditional. The MT-Bench paper found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans. The same paper names position bias, verbosity bias and self-enhancement bias, and notes limited reasoning ability. Those findings are about pairwise preference on chat answers; they do not transfer automatically to your domain, your rubric or your judge model.

  1. Label a sample by hand. Take a few dozen cases that span good, bad and borderline outputs, and have a person score them with the same rubric the judge gets.
  2. Measure agreement beyond chance. Percent agreement flatters you when one label dominates. Cohen's kappa corrects for chance agreement (kappa equals observed agreement minus expected agreement, divided by one minus expected agreement) and runs from -1 to 1. Its own caveat applies: the value depends on prevalence and bias, so no single threshold is universal. Look at the confusion matrix, not only the number.
  3. Fix the judge's inputs. Use a narrow rubric, binary or small-scale labels, and ask for the reasoning before the score. Pin the exact judge model version, because a silent update changes your trend line.
  4. Counter known biases. Randomise answer order in pairwise comparisons, avoid using the same model to generate and to judge when you can, and check whether scores correlate with answer length.
  5. Re-calibrate on change. New judge model, new rubric, new domain: repeat the comparison. Keep the labelled sample as a permanent calibration set.

Judge calls cost tokens on every run, so the cost and latency tactics from LLM cost, latency and prompt caching apply to your eval suite too. A smaller judge model is acceptable if, and only if, the calibration says so.

Tools: Ragas, DeepEval, TruLens and Phoenix

You can write all of this yourself in a few hundred lines, and for the deterministic retrieval metrics I often do. For LLM-judged metrics and for tracing, the established tools save real time. This is what their own documentation says, checked on 2 October 2026:

ToolWhat it isRAG metrics and workflow
RagasA metrics framework from the RAGAS paperContext precision, context recall, noise sensitivity, response relevancy, faithfulness; LLM-based and non-LLM variants; custom metrics supported
DeepEvalA test-style evaluation frameworkContextual relevancy, precision and recall for the retriever; answer relevancy and faithfulness for the generator; native pytest integration through the deepeval test run command for CI/CD
TruLensOpen-source evaluation and tracing, OpenTelemetry-native, maintained by SnowflakeThe RAG triad of context relevance, groundedness and answer relevance; works with LangChain, LlamaIndex and LangGraph
Arize PhoenixOpen-source AI observability and evaluation, built on OpenTelemetry and OpenInferenceTracing, evaluations with LLM, code or human labels, datasets and experiments to compare changes on the same inputs; Arize AX is the managed option

My practical advice: choose by where evals should live. If you want them as tests in the repository, DeepEval or a thin Ragas wrapper fits. If you want to inspect real production traces and attach scores to them, TruLens or Phoenix fit. Whatever you choose, keep the golden set and the labels in your own repository in a plain format, so switching tools costs an afternoon rather than a quarter. And remember that metrics with the same name can be computed differently across tools, so never compare scores between them.

Run evals in CI

An eval that nobody runs is documentation. The goal is that a change to chunking, the embedding model, the prompt or the generator cannot merge without producing numbers. I use two tiers, because the cost profiles differ.

  • Every pull request: retrieval only. Run recall@k, MRR and nDCG against a fixed index snapshot or fixture. These can be deterministic and need no LLM call, so they are fast, free and stable.
  • On a schedule or on relevant changes: the LLM-judged suite. Run faithfulness, answer relevance and correctness when the index, prompt, model or judge changes, and nightly otherwise. DeepEval's documentation describes a test command for running evals in CI/CD pipelines.
  • Gate on regression, not on a magic number. Compare to the last accepted baseline per stratum and fail on a drop beyond the noise you measured by running the suite several times.
  • Control the noise. Pin judge and generator versions, set temperature low, cache judge calls for unchanged inputs, and store every run's scores and retrieved chunk IDs so a failure can be diagnosed without re-running.
  • Close the loop. A failing production case becomes a golden-set row in the same pull request that fixes it.

For how this fits the broader practice of testing LLM features, including online monitoring and human review, see evals for LLM product features.

What I would do first

  1. Log, for every request, the query, the retrieved chunk IDs with scores and the final answer.
  2. Collect 30 to 50 real queries across your strata and label relevant chunks and a short reference answer.
  3. Compute recall@k, MRR and nDCG today. Fix retrieval before touching the prompt.
  4. Add faithfulness and answer relevance with an LLM judge, and calibrate it on a hand-labelled sample using kappa.
  5. Wire the cheap metrics into pull requests and the judged suite into a nightly job.
  6. Add every production failure to the golden set, with its diagnosis from the decision chain.

None of this needs a platform. It needs a labelled set, a few honest numbers and the habit of asking, for each bad answer, which half failed.

Sources

  1. Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)
  2. Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)
  3. Ragas documentation: available metrics
  4. Ragas documentation: faithfulness
  5. Ragas documentation: context precision
  6. Ragas documentation: context recall
  7. DeepEval documentation: metrics introduction
  8. TruLens
  9. Arize Phoenix documentation
  10. Wikipedia: Discounted cumulative gain
  11. Wikipedia: Mean reciprocal rank
  12. Wikipedia: Cohen's kappa

Frequently asked questions

How do you evaluate a RAG system?

Evaluate retrieval and generation separately on the same golden set of real queries. For retrieval, measure whether the right chunks were found and how high they ranked (recall@k, MRR, nDCG). For generation, measure whether the answer is supported by the retrieved context (faithfulness) and whether it addresses the question (answer relevance). Then use the two sets of scores to find which half failed.

What is the difference between recall@k, MRR and nDCG?

Recall@k asks whether the relevant material appears anywhere in the top k results. MRR is the average of 1 divided by the rank of the first relevant result, so it only cares about the first hit. nDCG scores the whole ranking with graded relevance, discounting results that appear lower, and is normalised to a range of 0 to 1.

What is faithfulness in RAG evaluation?

Faithfulness measures whether the claims in the generated answer are supported by the retrieved context. In Ragas it is the number of supported claims divided by the total number of claims in the response, a score between 0 and 1. A faithful answer can still be wrong if retrieval brought back the wrong context.

Can I trust an LLM as a judge?

Partly. The MT-Bench study found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans, but it also documented position, verbosity and self-enhancement bias. Treat the judge as an instrument: label a sample by hand, measure agreement, and recalibrate when the judge model or the rubric changes.

Which RAG evaluation tool should I use: Ragas, DeepEval, TruLens or Phoenix?

They overlap on metrics and differ in workflow. Ragas is a metrics framework, DeepEval is built around pytest-style tests and a CI command, TruLens centres on the RAG triad with tracing, and Arize Phoenix combines OpenTelemetry-based tracing with evaluations, datasets and experiments. I would pick by where you want the evals to live, not by metric names.

How do I know whether retrieval or generation caused a wrong answer?

Look at the retrieved chunks for that query. If the needed fact was never in the corpus, it is a content gap. If it was in the corpus but not in the top k, retrieval failed. If it was in the context but the answer ignores or contradicts it, generation failed, which shows up as low faithfulness or low answer relevance.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.