Tools/LLMOps & evals

Ragas review: the default RAG evaluation framework

Ragas scores RAG and agent pipelines with faithfulness, context precision and context recall, generates test data and records experiments. Apache-2.0, free library, judge model billed separately.

Type
Evaluation framework
Pricing
Apache-2.0 · paid platform

··10 min read

  • RAG evaluation
  • LLM-as-judge
  • Test sets
  • CI quality
  • Ragas
Diagram of a Ragas experiment loop: a frozen test set feeds a run of the application, metrics call a judge model per row, scores arrive with written reasons, and the result is compared with the baseline.

Key takeaways

  • Ragas is an Apache-2.0 Python library at version 0.4.3 whose metrics, faithfulness, context precision and context recall, have become the standard vocabulary of RAG evaluation.
  • Most metrics call a judge model, so a run costs money and inherits the judge's blind spots; pin the model version before comparing two runs.
  • The library stores no traces and draws no dashboards, so teams that need production observability pair it with a tool such as Langfuse or build the store themselves.
  • The hosted platform has no published price list, which makes the judge model, not the licence, the number to budget for.
  • Reference-free metrics measure whether an answer is supported, not whether it is correct; keep a small labelled set for correctness.

Ragas is an Apache-2.0 Python library for scoring LLM applications, and it has quietly become the vocabulary RAG evaluation is argued in: faithfulness, context precision and context recall all come out of its metric set. Its worth as a first evaluation tool for a retrieval pipeline is high; its worth as a production monitoring system is close to zero, because it computes numbers over a dataset you hand it and leaves storage, dashboards, sampling and alerting to somebody else.

It sits between the retrieval stack and the build: after LlamaIndex, LangChain or a hand-written retriever produces answers, Ragas scores them and the numbers end up next to the commit. It competes with DeepEval, TruLens and Phoenix, and with the evaluation features inside Langfuse, LangSmith and Braintrust; the difference is that Ragas stays a library, with no service to host and no vendor holding your test data.

What it is

The package installs from PyPI as pip install ragas and runs inside the process; version 0.4.3 was published on 13 January 2026 and requires Python 3.9 or newer. Metrics come in two kinds: LLM-based, where a judge model reads a row and returns a value with a written reason, and traditional, where strings are compared directly, such as BLEU, ROUGE, exact match and semantic similarity. Around that sits an experiments workflow, dataset, run, record, compare, which is what turns a pile of metric functions into a repeatable evaluation.

  • Licence and ownership: Apache-2.0, maintained by VibrantLabs; the repository shows about 16,000 stars and 1,700 forks.
  • RAG metrics: context precision, context recall, context entities recall, noise sensitivity, response relevancy and faithfulness, plus multimodal variants of faithfulness and relevance.
  • Agent metrics: tool call accuracy, tool call F1, agent goal accuracy and topic adherence.
  • Grounding and comparison: factual correctness, semantic similarity, BLEU, CHRF, ROUGE, string presence and exact match.
  • SQL and summarisation: execution-based DataCompy scoring, SQL query equivalence and a summarisation score.
  • Test data and scaffolding: synthetic test-set generation from your own documents, and a ragas quickstart rag_eval template that writes a runnable project.
  • Output: experiment results as CSV files in your repository, diffable like any other build artefact.

How it works

A run walks the rows one at a time. Each row carries the user input, the response the application produced and whatever context the retriever supplied. A metric takes that row, makes one or more LLM calls with a prompt that spells out the criterion, and returns a value together with the judge's reason; traditional metrics skip the call and compare strings. Rows are scored independently, so the set parallelises cleanly and any single row can be repeated when the average looks wrong.

How a Ragas experiment runsA five-stage loop. A frozen test set feeds a run of the application, metrics call a judge model over each row, the scores arrive with written reasons, and the experiment is compared with the baseline before the next build runs the same rows again.one row at a timetest setqueries, labelsrun appRAG pipelinemetricsjudge callsscoresvalue + reasoncomparevs baselinenext experiment: same rows, new build
The loop only means anything if the rows and the judge stay fixed: change either, and two experiments are not comparable.

The reason this style of evaluation spread is that the output is readable. A low faithfulness score arrives with the sentence the judge wrote about which claim was unsupported, which is a diff a human can act on rather than a number to argue about.

Getting started

The smallest useful test is one metric on one row, which is also how the README introduces the library.

import asyncio
from ragas.metrics.collections import AspectCritic
from ragas.llms import llm_factory

# The judge: llm_factory() takes OpenAI by default, Anthropic and others as options
llm = llm_factory("gpt-4o")

metric = AspectCritic(
    name="summary_accuracy",
    definition="Verify if the summary is accurate and captures key information.",
    llm=llm,
)

row = {
    "user_input": "summarise given text: the company reported an 8% rise in Q3 2024.",
    "response": "The company saw an 8% increase in Q3 2024.",
}

async def main():
    score = await metric.ascore(user_input=row["user_input"], response=row["response"])
    print(score.value, score.reason)

asyncio.run(main())

For a whole project the scaffolding is faster: uvx ragas quickstart rag_eval writes a package with rag.py, evals.py, a datasets folder and an experiments folder, uv sync installs it, and uv run python evals.py loads the dataset, queries the application, scores every row and appends the result to CSV. The docs default the template to OpenAI and show the same three lines switched to Anthropic, Google Gemini, Ollama or any OpenAI-compatible endpoint.

Performance and cost

Ragas adds no latency to the product; it adds cost and wall time to the build. The bill is rows times metrics times judge calls per row, and the expensive metrics are the ones that decompose an answer into claims and check each one separately.

MetricJudge work per rowNeeds a reference answerWhat a low score means
FaithfulnessSplits the answer into claims and checks each against the contextNoThe answer says more than the retrieved text supports
Context precisionRanks the retrieved chunks against the questionOptionalIrrelevant chunks are being pushed into the prompt
Context recallCompares retrieved context with the reference answerYesThe retriever never saw part of what the answer needs
Tool call accuracyCompares the agent's tool call and arguments with the expected oneYesThe agent picks the wrong tool or passes broken arguments

A 500-row set scored with four LLM metrics means thousands of judge calls per run. Teams keep it affordable by caching where the library allows it, by using a small judge model for the cheap metrics, and by running the full set only when a change is about to merge.

  • Score the cheap metrics on every push; run faithfulness and the agent metrics on merge or on a schedule.
  • Keep the dataset small and adversarial rather than large and random: 200 rows that broke before beat 5,000 rows that never fail.
  • Generate synthetic rows for coverage, then freeze them; regenerating the test set every cycle makes every score look like noise.

None of this is Ragas-specific: any LLM-as-judge setup bills by call and inherits the judge's blind spots. What Ragas decides is how much of that machinery you get for free, and how much you are left to build around it.

Pricing

The library is free and Apache-2.0, with no seat count and no usage cap. The commercial side, the hosted platform and direct support, is sold by the maintainers through contact and office hours, and no price list is published for it. The cost that therefore matters is the judge model: every LLM metric is metered by whichever provider you point it at, and synthetic test-set generation runs on the same meter.

  • Library: free, Apache-2.0, installed from PyPI, no account required.
  • Judge model: billed per call by your provider, typically the dominant line item of an evaluation budget.
  • Synthetic test data: also LLM work, billed the same way as scoring.
  • Hosted platform and support: sold direct, no public price; ask what is included before committing.

Where it shingles

Ragas scores rows, not systems. It will not sample production traffic, keep traces or tell you that p95 latency moved, and it cannot tell you whether the test questions are representative, which is the failure that makes a green pipeline useless. The release cadence is fast, 0.3.0 in July 2025, 0.4.0 in December 2025, 0.4.3 in January 2026, so versions have to be pinned. The repository carried 381 open issues at the time of writing, which is the usual price of being the default, and the docs for stable have trailed the released API.

ToolWhat it isWhere it winsWhat you give up
RagasMetric library with synthetic test dataThe metric vocabulary and a runnable project from one commandNo trace store, no dashboard, no production sampling
DeepEvalPytest-style metrics with a hosted platform behind itUnit-test ergonomics and a long metric listThe platform part is a separate commercial product
LangfuseOpen-source tracing with datasets and evaluationsTraces, prompts and scores in one placeRagas-style metric maths is one feature among many
Arize PhoenixOpen-source tracing and evaluation for LLM applicationsSpan-level debugging next to the scoresA narrower reference-free RAG metric set

The real choice is where the numbers live. Ragas is the right engine if something else stores the runs, a CSV in the repository, a table in your warehouse, an observability tool that accepts imported scores. It is the wrong purchase for a team that wants to open a URL and see whether quality regressed.

Verdict

Ragas earns its default status for retrieval work: the metrics are well defined, the output carries reasons, and a project is one command away. It should be taken as a scoring engine inside a workflow you already own, not as an evaluation platform.

  1. Use it if you have a retrieval pipeline and no golden set yet: synthetic test-set generation plus reference-free metrics gives a first measurement in a day.
  2. Use it if evaluation has to run in CI, where a CSV diff per pull request is exactly the right amount of signal.
  3. Do not pick it as the only evaluation tool for an agent product: agent metrics exist, but trajectories, tool side effects and cost per task need a harness that stores runs.
  4. Do not pick it when an answer has to be justified to a regulator or a customer from stored evidence; nothing is persisted unless you persist it.
  5. Budget the judge model before the licence: a pinned small model for the cheap metrics and a capable one for faithfulness is the usual split.
Ragas is the cheapest way to stop arguing about whether a change helped: pin the judge, freeze the rows, read the diff. Everything it does not do is work you still have to pay for.

Sources

  1. Ragas documentation: introduction
  2. Ragas documentation: quick start
  3. Ragas documentation: list of available metrics
  4. Ragas on GitHub
  5. Ragas on PyPI
  6. Ragas website

Frequently asked questions

Is Ragas free to use?

The Python library is Apache-2.0 and free, including metrics, synthetic test-set generation and the experiments workflow. The hosted platform and enterprise support are sold directly by VibrantLabs, and no prices are published for them.

Which Ragas metrics need a golden set?

Context recall and factual correctness compare against a reference answer, while faithfulness, response relevancy and the aspect critic work from the question, the context and the response alone. Most teams start reference-free and add labels only for the rows they keep getting wrong.

Does Ragas replace an observability tool?

No. It scores a dataset you provide and writes results to CSV; it does not collect traces, sample production traffic or render dashboards. Use it as the metric layer under something that stores runs.

How much does a Ragas evaluation cost to run?

Whatever the judge model charges. Faithfulness issues several LLM calls per row because it checks claim by claim, so a 500-row set across four metrics runs to thousands of calls; caching and a smaller judge model are the levers.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.