> How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.
>
> Web page: https://balazscsorba.com/blog/rag-evaluation-metrics · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/rag-evaluation-metrics.md) · [Magyar](https://balazscsorba.com/hu/blog/rag-evaluation-metrics.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: RAG evaluation, how to evaluate RAG, RAG evaluation metrics, recall@k MRR nDCG, faithfulness vs answer relevance, golden dataset for RAG, LLM as a judge calibration, Ragas vs DeepEval vs TruLens, RAG evals in CI, retrieval vs generation failure

[Blog](https://balazscsorba.com/blog)/RAG & retrieval

# Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed

How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·12 min read

-   RAG evaluation
-   Retrieval metrics
-   LLM-as-judge
-   Golden set
-   Ragas

![Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.](https://balazscsorba.com/images/blog/rag-evaluation-metrics/cover.webp?v=5cf36c4462)

## Key takeaways

-   A RAG system has two halves that fail differently, so score them separately: retrieval with recall@k, MRR and nDCG, generation with faithfulness and answer relevance.
-   Recall@k is the ceiling. If the needed chunk is not in the top k, no prompt or model can produce a grounded answer, so check retrieval first when an answer is wrong.
-   A golden set of real user queries beats synthetic questions. Start small, label what a correct answer needs, and grow it from every production failure.
-   An LLM judge is a measuring instrument, not a ground truth. Calibrate it against human labels on a sample, report chance-corrected agreement and re-check it whenever the judge model changes.
-   Run cheap retrieval checks on every pull request and the LLM-judged suite on a schedule or when the index, prompt or model changes, and gate on regression against a baseline.

On this page

1.  [Two halves, seven numbers](https://balazscsorba.com/#two-halves)
2.  [Retrieval metrics: recall@k, MRR and nDCG](https://balazscsorba.com/#retrieval-metrics)
3.  [Generation metrics: faithfulness and answer relevance](https://balazscsorba.com/#generation-metrics)
4.  [Build the golden set from real queries](https://balazscsorba.com/#golden-set)
5.  [Diagnose: did retrieval or generation fail?](https://balazscsorba.com/#diagnosis)
6.  [LLM-as-judge: calibrate it against humans](https://balazscsorba.com/#llm-judge)
7.  [Tools: Ragas, DeepEval, TruLens and Phoenix](https://balazscsorba.com/#tools)
8.  [Run evals in CI](https://balazscsorba.com/#ci)
9.  [What I would do first](https://balazscsorba.com/#first-steps)
10.  [Sources](https://balazscsorba.com/#sources)

Most RAG systems are shipped on the strength of a demo. Someone asks five questions, the answers look right, and the team moves on. Then the corpus grows, the chunking changes, the model is upgraded, and nobody can say whether the system got better or worse. The honest answer to "is it good?" is a number you can re-compute, and for RAG that takes more than one number.

The reason is structural. A RAG system is two systems in a row: a retriever that decides what the model sees, and a generator that decides what to say about it. Each fails in its own way, needs its own metrics and gets fixed by different work. If you only score the final answer, a wrong answer tells you that something broke, but not where.

This article is how I would set up RAG evaluation for a product team: the metrics and what each one actually measures, a golden set built from real queries, an LLM judge you have checked against humans, the tools worth looking at, evals that run in CI, and a short procedure for diagnosing whether retrieval or generation failed. It builds on my posts about [evals for LLM product features](https://balazscsorba.com/blog/llm-evals-for-product-features) and the [RAG pipeline itself](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking).

## Two halves, seven numbers

The [RAGAS paper](https://arxiv.org/abs/2309.15217) frames RAG evaluation along three dimensions: the retrieval system's ability to find relevant and focused passages, the generator's ability to use those passages faithfully, and the quality of the output. In practice I split the numbers by the question they answer. Retrieval metrics ask _did the right material reach the model, and in a good order?_ Generation metrics ask _given what it saw, did the model behave?_

The table lists the metrics I actually use. Names differ between tools, so I describe what each one measures rather than what a library calls it.

Metric

Layer

The question it answers

Needs labels?

recall@k

Retrieval

Is the relevant material anywhere in the top k results?

Relevant chunks per query

MRR

Retrieval

How high is the first relevant result? The mean of 1 divided by its rank.

Relevant chunks per query

nDCG

Retrieval

Is the whole ranking in a good order, with graded relevance and a discount for lower positions?

Graded relevance

Context precision

Retrieval (ranking)

Are relevant chunks ranked above irrelevant ones in what the model received?

Optional reference answer in Ragas

Context recall

Retrieval

Is everything the reference answer says supported by the retrieved context?

Reference answer

Faithfulness

Generation

Is each claim in the answer supported by the retrieved context?

No

Answer relevance

Generation

Does the answer address the question that was asked?

No

Note that context precision sits with retrieval, even though it is often listed beside the generation metrics. It is computed over the retrieved chunks, so it tells you about the ranker, not about the model's writing. Keeping the layer column honest is what makes the diagnosis later in this article possible.

## Retrieval metrics: recall@k, MRR and nDCG

Retrieval metrics come from search evaluation and need one thing the generation metrics do not: a notion of which chunks are relevant for each query.

-   **Recall@k** asks whether the relevant material is in the top k. It is the most important retrieval number, because it is a ceiling. If the chunk that contains the answer is not among the k chunks you pass to the model, the model cannot ground its answer in it, whatever the prompt says. Choose k to match what you actually send to the generator.
-   **MRR** (mean reciprocal rank) averages 1 divided by the rank of the first relevant result: 1 for first place, 0.5 for second. It rewards putting a good chunk at the top. Its documented limitation is that only the first relevant result counts and any further ones are ignored, so it suits questions with one right passage and misleads on questions that need several.
-   **nDCG** sums graded relevance scores with a logarithmic discount for lower positions, then divides by the ideal ordering so the result lies between 0 and 1. It is the metric to use when relevance is not binary, for example a chunk that fully answers versus one that only mentions the topic, and when ordering matters because the context window or a reranker's cut-off truncates the list.

I read them together. High recall@k with a poor MRR or nDCG means the right chunk is retrieved but buried, which a reranker usually fixes (see the [pipeline post](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking) for chunking, hybrid search and reranking). Low recall@k means the right chunk is never found, which is a problem with chunking, embeddings, query rewriting or the index, and no reranker can repair a candidate list that does not contain the answer.

These metrics need relevance labels, which is the expensive part. The Ragas documentation describes an ID-based and a non-LLM variant of context recall that compare retrieved chunk IDs or text to reference contexts without an LLM call. Those are cheap, deterministic and perfect for CI, if your golden set stores which chunks are relevant.

## Generation metrics: faithfulness and answer relevance

Generation metrics judge the answer given what the model was shown. **Faithfulness** is the central one. In the Ragas definition it is the number of claims in the response that the retrieved context supports, divided by the total number of claims: the response is decomposed into statements, each statement is checked against the context, and the ratio is the score, from 0 to 1. It needs no reference answer, which is why it is so useful on production traffic.

**Answer relevance** asks whether the answer addresses the question at all. An answer can be perfectly faithful to the context and still dodge the question, for example by summarising a document when the user asked for a number. TruLens calls the equivalent triad _context relevance, groundedness and answer relevance_; Ragas calls it response relevancy; DeepEval calls it answer relevancy. Same idea, different labels.

Two properties are easy to miss. First, faithfulness measures agreement with the retrieved context, not with the truth. If retrieval returned a stale policy document, a perfectly faithful answer is wrong. Second, a high score can be earned by saying little: a model that refuses or hedges makes few claims, and few claims can all be supported. Always read faithfulness next to answer relevance and a correctness check against the golden answer.

**Faithful is not correct**

Faithfulness only tells you the answer matches the chunks. Whether the chunks were the right ones is a retrieval question, and whether the final answer matches the golden answer is a separate correctness score. I track all three, because the combinations are what point at the cause.

## Build the golden set from real queries

Everything above depends on a golden set: queries with the expectations you score against. Synthetic questions generated from your documents are a useful bootstrap, but they share the vocabulary of the documents, which real users do not. Real queries have typos, abbreviations, vague phrasing, product codes and questions the corpus cannot answer. I would build the set from them.

1.  **Sample real queries.** Take them from search logs, support tickets or chat transcripts. Remove personal data first, and mind your legal basis: see my notes on [GDPR and LLM APIs](https://balazscsorba.com/blog/gdpr-llm-api-eu-data-residency).
2.  **Stratify, do not just sample.** Include the common questions, the long tail, multi-document questions, questions with exact identifiers, and questions that should get "I don't know" because the corpus has no answer.
3.  **Label what a correct answer needs.** For each query, record the relevant chunks or documents (for recall@k, MRR and nDCG), a short reference answer, and where it matters the facts the answer must include.
4.  **Have a domain expert label it, and measure whether two people agree.** If two humans disagree about what is relevant, no metric can resolve that for you.
5.  **Version it and keep a held-out part.** Treat the set like code. Keep a slice you never tune against, so improvements are not just overfitting to the examples you stared at.
6.  **Feed it from production.** Every thumbs-down, escalation or wrong answer found in review becomes a new case, with its label. This is how the set stays representative as the product changes.

How big? I would start with a few dozen well-labelled queries and grow from there; a small clean set that you trust beats a large one nobody checked. The point is not statistical power at the start, it is to have something that makes a regression visible the day it happens. As the set grows, report results per stratum, because an average hides a collapse in one query type.

## Diagnose: did retrieval or generation fail?

When a wrong answer is reported, resist the urge to edit the prompt. Walk the chain from the bottom, using the stored retrieval results for that query, and stop at the first question you answer with "no".

The order matters: each question assumes the ones above it have passed. Most of the surprises I see are in the first two.

The branches map to different work. A **content gap** is an editorial or ingestion problem: the document is missing, outdated or was never parsed properly. A **retrieval failure** is fixed in chunking, embeddings, hybrid search, query rewriting or reranking. An **unfaithful** answer is a generation problem: tighten the instruction to answer only from the context, reduce irrelevant context that invites speculation, or change the model. An **off-target** answer is usually the prompt, or a retrieval problem in disguise where the context is on topic but not on the question.

The last box matters too. If every metric passes and the answer is still wrong, suspect the golden label, the reference answer or the judge itself before the system. For agentic setups, where the model decides what to retrieve and may retrieve several times, the same logic applies per retrieval step; see [RAG in 2026: hybrid, agentic and long context](https://balazscsorba.com/blog/rag-2026-hybrid-agentic-long-context).

## LLM-as-judge: calibrate it against humans

Faithfulness, answer relevance and most context metrics are scored by an LLM. Ragas notes that its LLM-based metrics may use one or more LLM calls to produce a score. That makes the judge part of your measurement setup, and an instrument needs calibration.

The evidence is encouraging but not unconditional. The [MT-Bench paper](https://arxiv.org/abs/2306.05685) found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans. The same paper names position bias, verbosity bias and self-enhancement bias, and notes limited reasoning ability. Those findings are about pairwise preference on chat answers; they do not transfer automatically to your domain, your rubric or your judge model.

1.  **Label a sample by hand.** Take a few dozen cases that span good, bad and borderline outputs, and have a person score them with the same rubric the judge gets.
2.  **Measure agreement beyond chance.** Percent agreement flatters you when one label dominates. Cohen's kappa corrects for chance agreement (kappa equals observed agreement minus expected agreement, divided by one minus expected agreement) and runs from -1 to 1. Its own caveat applies: the value depends on prevalence and bias, so no single threshold is universal. Look at the confusion matrix, not only the number.
3.  **Fix the judge's inputs.** Use a narrow rubric, binary or small-scale labels, and ask for the reasoning before the score. Pin the exact judge model version, because a silent update changes your trend line.
4.  **Counter known biases.** Randomise answer order in pairwise comparisons, avoid using the same model to generate and to judge when you can, and check whether scores correlate with answer length.
5.  **Re-calibrate on change.** New judge model, new rubric, new domain: repeat the comparison. Keep the labelled sample as a permanent calibration set.

Judge calls cost tokens on every run, so the cost and latency tactics from [LLM cost, latency and prompt caching](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing) apply to your eval suite too. A smaller judge model is acceptable if, and only if, the calibration says so.

## Tools: Ragas, DeepEval, TruLens and Phoenix

You can write all of this yourself in a few hundred lines, and for the deterministic retrieval metrics I often do. For LLM-judged metrics and for tracing, the established tools save real time. This is what their own documentation says, checked on 2 October 2026:

Tool

What it is

RAG metrics and workflow

Ragas

A metrics framework from the RAGAS paper

Context precision, context recall, noise sensitivity, response relevancy, faithfulness; LLM-based and non-LLM variants; custom metrics supported

DeepEval

A test-style evaluation framework

Contextual relevancy, precision and recall for the retriever; answer relevancy and faithfulness for the generator; native pytest integration through the deepeval test run command for CI/CD

TruLens

Open-source evaluation and tracing, OpenTelemetry-native, maintained by Snowflake

The RAG triad of context relevance, groundedness and answer relevance; works with LangChain, LlamaIndex and LangGraph

Arize Phoenix

Open-source AI observability and evaluation, built on OpenTelemetry and OpenInference

Tracing, evaluations with LLM, code or human labels, datasets and experiments to compare changes on the same inputs; Arize AX is the managed option

My practical advice: choose by where evals should live. If you want them as tests in the repository, DeepEval or a thin Ragas wrapper fits. If you want to inspect real production traces and attach scores to them, TruLens or Phoenix fit. Whatever you choose, keep the golden set and the labels in your own repository in a plain format, so switching tools costs an afternoon rather than a quarter. And remember that metrics with the same name can be computed differently across tools, so never compare scores between them.

## Run evals in CI

An eval that nobody runs is documentation. The goal is that a change to chunking, the embedding model, the prompt or the generator cannot merge without producing numbers. I use two tiers, because the cost profiles differ.

-   **Every pull request: retrieval only.** Run recall@k, MRR and nDCG against a fixed index snapshot or fixture. These can be deterministic and need no LLM call, so they are fast, free and stable.
-   **On a schedule or on relevant changes: the LLM-judged suite.** Run faithfulness, answer relevance and correctness when the index, prompt, model or judge changes, and nightly otherwise. DeepEval's documentation describes a test command for running evals in CI/CD pipelines.
-   **Gate on regression, not on a magic number.** Compare to the last accepted baseline per stratum and fail on a drop beyond the noise you measured by running the suite several times.
-   **Control the noise.** Pin judge and generator versions, set temperature low, cache judge calls for unchanged inputs, and store every run's scores and retrieved chunk IDs so a failure can be diagnosed without re-running.
-   **Close the loop.** A failing production case becomes a golden-set row in the same pull request that fixes it.

For how this fits the broader practice of testing LLM features, including online monitoring and human review, see [evals for LLM product features](https://balazscsorba.com/blog/llm-evals-for-product-features).

## What I would do first

1.  Log, for every request, the query, the retrieved chunk IDs with scores and the final answer.
2.  Collect 30 to 50 real queries across your strata and label relevant chunks and a short reference answer.
3.  Compute recall@k, MRR and nDCG today. Fix retrieval before touching the prompt.
4.  Add faithfulness and answer relevance with an LLM judge, and calibrate it on a hand-labelled sample using kappa.
5.  Wire the cheap metrics into pull requests and the judged suite into a nightly job.
6.  Add every production failure to the golden set, with its diagnosis from the decision chain.

None of this needs a platform. It needs a labelled set, a few honest numbers and the habit of asking, for each bad answer, which half failed.

## Sources

1.  [Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)](https://arxiv.org/abs/2309.15217)
2.  [Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)](https://arxiv.org/abs/2306.05685)
3.  [Ragas documentation: available metrics](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/)
4.  [Ragas documentation: faithfulness](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/faithfulness/)
5.  [Ragas documentation: context precision](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_precision/)
6.  [Ragas documentation: context recall](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_recall/)
7.  [DeepEval documentation: metrics introduction](https://deepeval.com/docs/metrics-introduction)
8.  [TruLens](https://www.trulens.org/)
9.  [Arize Phoenix documentation](https://arize.com/docs/phoenix)
10.  [Wikipedia: Discounted cumulative gain](https://en.wikipedia.org/wiki/Discounted_cumulative_gain)
11.  [Wikipedia: Mean reciprocal rank](https://en.wikipedia.org/wiki/Mean_reciprocal_rank)
12.  [Wikipedia: Cohen's kappa](https://en.wikipedia.org/wiki/Cohen%27s_kappa)

## Frequently asked questions

How do you evaluate a RAG system?

Evaluate retrieval and generation separately on the same golden set of real queries. For retrieval, measure whether the right chunks were found and how high they ranked (recall@k, MRR, nDCG). For generation, measure whether the answer is supported by the retrieved context (faithfulness) and whether it addresses the question (answer relevance). Then use the two sets of scores to find which half failed.

What is the difference between recall@k, MRR and nDCG?

Recall@k asks whether the relevant material appears anywhere in the top k results. MRR is the average of 1 divided by the rank of the first relevant result, so it only cares about the first hit. nDCG scores the whole ranking with graded relevance, discounting results that appear lower, and is normalised to a range of 0 to 1.

What is faithfulness in RAG evaluation?

Faithfulness measures whether the claims in the generated answer are supported by the retrieved context. In Ragas it is the number of supported claims divided by the total number of claims in the response, a score between 0 and 1. A faithful answer can still be wrong if retrieval brought back the wrong context.

Can I trust an LLM as a judge?

Partly. The MT-Bench study found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans, but it also documented position, verbosity and self-enhancement bias. Treat the judge as an instrument: label a sample by hand, measure agreement, and recalibrate when the judge model or the rubric changes.

Which RAG evaluation tool should I use: Ragas, DeepEval, TruLens or Phoenix?

They overlap on metrics and differ in workflow. Ragas is a metrics framework, DeepEval is built around pytest-style tests and a CI command, TruLens centres on the RAG triad with tracing, and Arize Phoenix combines OpenTelemetry-based tracing with evaluations, datasets and experiments. I would pick by where you want the evals to live, not by metric names.

How do I know whether retrieval or generation caused a wrong answer?

Look at the retrieved chunks for that query. If the needed fact was never in the corpus, it is a content gap. If it was in the corpus but not in the top k, retrieval failed. If it was in the context but the answer ignores or contradicts it, generation failed, which shows up as low faithfulness or low answer relevance.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [Reducing LLM hallucinations in production: grounding, citations and knowing when to say no](https://balazscsorba.com/blog/llm-hallucination-grounding-citations)
-   [pgvector or a vector database? How to choose vector storage in 2026](https://balazscsorba.com/blog/pgvector-vs-vector-databases)
-   [GraphRAG and knowledge-graph RAG: when a graph beats vector search](https://balazscsorba.com/blog/graphrag-knowledge-graph-rag)
-   [Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure](https://balazscsorba.com/blog/semantic-product-search-b2b)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
