Tools/LLMOps & evals
DeepEval review: pytest for LLM outputs, and the judge bill
DeepEval runs LLM checks as pytest-style tests, with built-in judge metrics. The library is free and Apache-2.0; the judge calls and the data flow are the real cost.
- Type
- Evaluation framework
- Pricing
- Apache-2.0 · free; Confident AI from $0, Starter $200 a month
Balázs Csorba··8 min read
- LLM evaluation
- pytest
- LLM-as-a-judge
- Red teaming
- Open source

Key takeaways
- DeepEval is an Apache-2.0 Python framework that runs LLM checks as unit tests, and the library works with no account.
- Most predefined metrics use another model as the judge, so every run can cost judge tokens and can return a different score.
- With default settings the judge is OpenAI, so test cases leave your network. A local judge such as Ollama keeps them inside, at a quality you have to measure.
- Confident AI's cloud has a free plan, Starter at $200 a month and Team at $2,000 a month, with a DPA for every customer and an EU region on every plan.
- Choose Ragas, Promptfoo or Braintrust for an Apache-2.0 metrics toolkit, a declarative CLI with red teaming, or a hosted UI with its own meter.
DeepEval is an open-source Python framework that runs LLM checks as unit tests, with more than thirty built-in metrics and an optional cloud platform called Confident AI. The verdict up front: take it if your engineers write Python, already run tests in CI and want model behaviour checked on every pull request. Skip it if the default OpenAI judge is not acceptable for your data, or if non-engineers need a hosted workspace before anything else.
It sits next to Ragas, Promptfoo and Braintrust, which I have reviewed separately. For the process around the tests, from reading traces by hand to a gate in CI, see LLM evals for product features.
What it is
DeepEval is the Apache-2.0 library behind Confident AI, which describes itself as the enterprise AI evals and observability platform. The library runs on your machine and needs no account. The current release on PyPI is 4.2.8, published on 2 October 2026, and it needs Python 3.9 or newer. Install it with pip install -U deepeval.
- Metrics. More than thirty named metrics in nine groups, from G-Eval and agent metrics to retrieval, multi-turn, safety and image checks.
- Judges. Almost all predefined metrics use an LLM as the judge, and you can point them at OpenAI, Azure OpenAI, Anthropic, Gemini, Ollama or LiteLLM.
- Red teaming. A separate Apache-2.0 package, DeepTeam, covers attacks and vulnerabilities.
How it works
A test is ordinary Python. Each metric receives a test case, which holds the input, the actual output and, depending on the metric, the expected output, the retrieval context or the tools the agent called. For a judge-based metric, the test case and the criteria go to the judge model, which returns a score from 0 to 1 and a reason. The threshold turns the score into a pass or a fail, and assert_test fails the test when the score is below it.
The judge is the part that matters for cost and reliability. The metrics page says almost all predefined metrics use an LLM as the judge, and that any judge can be used, including OpenAI, Azure OpenAI, Ollama, Anthropic, Gemini and LiteLLM. You can also wrap your own model by subclassing DeepEvalBaseLLM. The FAQ says the judge defaults to OpenAI when you specify no model.
Getting started
Export an OpenAI key for the judge, install the package and write the test in a file such as test_example.py. The example checks an answer with G-Eval, where you describe the criteria in plain language and DeepEval writes the evaluation steps from them. Run it with deepeval test run test_example.py. The test fails if the correctness score is below 0.5.
# test_example.py
from deepeval import assert_test
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
def test_refund_answer():
correctness = GEval(
name='Correctness',
criteria='Is the actual output a correct answer to the input, consistent with the expected output?',
evaluation_params=[
SingleTurnParams.INPUT,
SingleTurnParams.ACTUAL_OUTPUT,
SingleTurnParams.EXPECTED_OUTPUT,
],
threshold=0.5,
)
test_case = LLMTestCase(
input='Can I return shoes after 30 days?',
actual_output='You can return them within 30 days of delivery, if they are unworn.',
expected_output='Returns are accepted within 30 days of delivery if the item is unworn.',
)
assert_test(test_case, [correctness])Metrics, and what the judge costs
Metric choice decides the bill. Faithfulness extracts the claims in an answer, checks each claim against the retrieval context, and scores the share of claims that do not contradict it, with a default threshold of 0.5. The judge has more to read when an answer makes many claims, so the cost grows with the answer. The evaluate() function brings caching, parallelisation, cost tracking and error handling, while a standalone measure() call loses those optimisations. The docs also warn that many metric calls at once can trigger rate-limit errors, so cap the concurrency to your provider's limit.
| Group | Named metrics |
|---|---|
| Custom | G-Eval, DAG, Arena G-Eval, JevEval, and your own code metrics such as BLEU or ROUGE |
| Agents, trajectory | Task Completion, Step Efficiency, Plan Adherence, Plan Quality |
| Agents, components | Tool Correctness, Argument Correctness |
| Retriever | Contextual Relevancy, Contextual Precision, Contextual Recall |
| Generator | Answer Relevancy, Faithfulness |
| Multi-turn chatbots | Knowledge Retention, Role Adherence, Conversation Completeness, Conversation Relevancy |
| Safety | Bias, Toxicity, Non-Advice, Misuse, PIILeakage, Role Violation |
| Image | Image Coherence, Image Helpfulness, Image Reference, Text-to-Image, Image-Editing |
| Other | Hallucination, JSON Correctness, Summarization, Ragas |
Goldens, red teaming and CI
Hand-written test cases run out quickly, so DeepEval can generate them. generate_goldens_from_docs takes a list of document paths, reads .txt, .docx, .pdf and Markdown files, stores the chunks in chromadb and has a critic model score each chunk from 0 to 1. By default each golden also gets an expected output. Without an OpenAI key you must supply your own embedding model and LLM, because the default embedder is text-embedding-3-small. The docs describe no review step, so have someone read a sample before a generated set becomes a regression suite.
Red teaming is a separate package. DeepTeam is an Apache-2.0 framework to red team LLMs and AI agents, with more than 50 ready-made vulnerabilities and more than 20 research-backed attack methods, in single-turn and multi-turn form. Its LLM-as-a-judge metrics run on your machine and return a pass or fail with reasoning. It maps its checks to the OWASP Top 10 for LLMs 2025, the OWASP Top 10 for Agents 2026, NIST AI RMF and MITRE ATLAS, and it installs with pip install -U deepteam. The platform's AI red-teaming module sits in the Enterprise ++ tier of the pricing page.
The CI pattern in the docs is short. Store the judge key as OPENAI_API_KEY, add CONFIDENT_API_KEY only if the results should reach the platform, and run deepeval test run. Without the key the same tests still run locally. Add --official (or -o) to mark a run as the official baseline on Confident AI.
- name: Run LLM tests
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
CONFIDENT_API_KEY: ${{ secrets.CONFIDENT_API_KEY }}
run: poetry run deepeval test run test_llm_app.pyCost and deployment
The library costs nothing to run beyond the judge. The platform has four plans, billed monthly per organisation, and the pricing page carries no 'as of' date, so check it again before you budget.
| Plan | Price | Included | What it adds |
|---|---|---|---|
| Free | $0 | 2 seats, 1 project, 1 GB-month of trace data | 5 test runs a week |
| Starter | $200 a month | Unlimited seats, 5 projects, 5 GB-months, then $1 per GB-month | Online evals, annotation queues, real-time alerts |
| Team | $2,000 a month | Unlimited seats and projects, 75 GB-months, then $1 per GB-month | SOC 2, SSO, custom roles, custom contracts and SLAs |
| Enterprise | Custom | Unlimited usage | On-prem, custom data residency, 24x7 support, red-teaming module (Enterprise ++) |
For comparison, Braintrust's Starter tier is free with 14 days of retention, and its Pro plan is $249 a month, with processed data at $3 per GB and scores at $1.50 per 1,000 beyond the included amounts. The Confident AI Starter plan is a flat $200 a month with no per-seat fee, and it includes online evals and annotation queues.
There are three separate data flows, and they answer different questions. The judge receives the test case, so with no model configured that content goes to OpenAI, from your laptop or CI runner. Confident AI receives results only when CONFIDENT_API_KEY is set, and the GitHub README says that when you use the cloud platform all test cases are logged automatically. Telemetry is the third flow: by default DeepEval records basic, non-identifying counts, such as how many evaluations ran and which metrics were used, and the data-privacy page names PostHog as the only destination. Set DEEPEVAL_TELEMETRY_OPT_OUT=1 to switch it off.
On the platform, the FAQ says data is stored in a private AWS cloud that only your organisation can access. The default region is the United States, and the EU is available on every plan. Choose it at login or sign-up, or run deepeval set-confident-region EU, which configures DeepEval only. The region page adds that an API key alone does not decide where the data goes, so set the two endpoints as well.
deepeval set-confident-region EU
export CONFIDENT_BASE_URL=https://eu.api.confident-ai.com
export CONFIDENT_OTEL_ENDPOINT=https://eu.otel.confident-ai.com/v1/tracesFor the contract, the subprocessor list, last modified on 18 February 2026, says every customer is given a data processing agreement. Personal data stays in the chosen region unless it has to move for performance or availability, or as agreed. The list names OpenAI for AI model inference, US only, and AWS, ClickHouse, Supabase and PostHog as US and Europe. I could not find a retention period for the cloud in any page I opened, so ask for it in writing before production data goes in. Self-hosting is an Enterprise option, and the on-premise setup points the same two variables at your own hosts.
Where it falls short
Three weaknesses matter most. First, scores are not repeatable. The G-Eval docs say the metric is not deterministic, so a score near the threshold will sometimes flip, and the judge caps the quality of everything you measure with it. Second, the surface is large. The metric list is long, the document route for generated goldens needs chromadb plus several LangChain packages, and the docs show a TypeScript command next to the Python one, although everything I checked for this review is Python. Third, the platform's price steps are coarse. The Free plan's five test runs a week will not carry a busy pipeline, and the jump from $200 to $2,000 a month is large for a team that only wants shared history.
Verdict
DeepEval is the right default for a Python team that wants LLM behaviour tests in the same runner and the same pull request as the rest of its suite. The library and DeepTeam are both Apache-2.0, the metric list covers agents, retrieval, conversations and safety without writing your own judges, and the local mode works with no account. I would not adopt it where the test data cannot reach a judge you trust, or where non-engineers need a hosted workspace from day one. Pick an alternative in those cases.
- Ragas if you want an Apache-2.0 toolkit with pre-built metrics and test data generation, and you do not need the Confident AI platform.
- Promptfoo if you want declarative configs that compare prompts and models from the command line in CI, with red teaming in the same tool. Its MIT repository now says Promptfoo is part of OpenAI, so weigh that against your vendor rules.
- Braintrust if you want a hosted UI and a click-through DPA on Pro, and you accept its meter, where usage above the included data and scores is billed on top of the $249 a month.
Sources
- DeepEval documentation: getting started
- GitHub: confident-ai/deepeval, the README and licence
- PyPI: deepeval, the latest release and Python requirement
- DeepEval documentation: metrics introduction
- DeepEval documentation: G-Eval
- DeepEval documentation: Faithfulness
- DeepEval documentation: Tool Correctness
- DeepEval documentation: generate goldens from documents
- DeepEval documentation: unit testing in CI/CD
- DeepEval FAQ
- DeepEval documentation: data privacy
- GitHub: confident-ai/deepteam, the red-teaming framework
- Confident AI pricing
- Confident AI documentation: data residency
- Confident AI subprocessor list
- GitHub: explodinggradients/ragas, the README
- GitHub: promptfoo/promptfoo, the README
- Braintrust pricing
Frequently asked questions
Does DeepEval send my data anywhere?
Yes, to the judge model you configure. With no model set, the FAQ says the judge defaults to OpenAI. Results reach Confident AI only when you set a key, and the default storage region is the United States, with the EU available on every plan.
How much does DeepEval cost?
The library is free under an Apache-2.0 licence. Confident AI has a free plan with two seats and five test runs a week, Starter at $200 a month and Team at $2,000 a month, billed monthly per organisation. Judge calls are extra and depend on the model you choose.
Are the scores repeatable?
Not exactly. The G-Eval docs say the metric is not deterministic, so give thresholds some margin and re-run a failing case before you treat it as a regression. Tool Correctness is different, because its core score is deterministic.
Can I use a local model as the judge?
Yes. The metrics page lists Ollama among the judges, and you can wrap any other model by subclassing DeepEvalBaseLLM. Check the judge against cases you have labelled yourself, because a weaker judge moves every score.