[{"data":1,"prerenderedAt":841},["ShallowReactive",2],{"tool-deepeval-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":35,"cover":90,"og":91,"expertise":92,"locales":93,"lang":94,"title":97,"description":98,"coverAlt":99,"url":25,"pricing":100,"kind":101,"metaTitle":102,"takeaways":103,"faq":109,"toc":122,"blocks":150,"others":500},"deepeval","2026-10-06",8,"llmops",[9,10,11,12,13],"LLM evaluation","pytest","LLM-as-a-judge","Red teaming","Open source",[4,15,16,17,18,19,20,21],"deepeval vs ragas","llm evaluation framework","pytest for llm outputs","g-eval metric","llm judge cost","deepeval pricing","llm red teaming open source",[23,26,29,32],{"name":24,"url":25},"DeepEval","https:\u002F\u002Fdeepeval.com",{"name":27,"url":28},"Confident AI","https:\u002F\u002Fwww.confident-ai.com",{"name":30,"url":31},"Pytest","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPytest",{"name":33,"url":34},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[36,39,42,45,48,51,54,57,60,63,66,69,72,75,78,81,84,87],{"title":37,"url":38},"DeepEval documentation: getting started","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fgetting-started",{"title":40,"url":41},"GitHub: confident-ai\u002Fdeepeval, the README and licence","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepeval",{"title":43,"url":44},"PyPI: deepeval, the latest release and Python requirement","https:\u002F\u002Fpypi.org\u002Fproject\u002Fdeepeval\u002F",{"title":46,"url":47},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":49,"url":50},"DeepEval documentation: G-Eval","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-llm-evals",{"title":52,"url":53},"DeepEval documentation: Faithfulness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-faithfulness",{"title":55,"url":56},"DeepEval documentation: Tool Correctness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-tool-correctness",{"title":58,"url":59},"DeepEval documentation: generate goldens from documents","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fsynthesizer-generate-from-docs",{"title":61,"url":62},"DeepEval documentation: unit testing in CI\u002FCD","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fevaluation-unit-testing-in-ci-cd",{"title":64,"url":65},"DeepEval FAQ","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Ffaq",{"title":67,"url":68},"DeepEval documentation: data privacy","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fdata-privacy",{"title":70,"url":71},"GitHub: confident-ai\u002Fdeepteam, the red-teaming framework","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam",{"title":73,"url":74},"Confident AI pricing","https:\u002F\u002Fwww.confident-ai.com\u002Fpricing",{"title":76,"url":77},"Confident AI documentation: data residency","https:\u002F\u002Fwww.confident-ai.com\u002Fdocs\u002Fsettings\u002Fdata-residency",{"title":79,"url":80},"Confident AI subprocessor list","https:\u002F\u002Fwww.confident-ai.com\u002Fsubprocessors-list",{"title":82,"url":83},"GitHub: explodinggradients\u002Fragas, the README","https:\u002F\u002Fgithub.com\u002Fexplodinggradients\u002Fragas",{"title":85,"url":86},"GitHub: promptfoo\u002Fpromptfoo, the README","https:\u002F\u002Fgithub.com\u002Fpromptfoo\u002Fpromptfoo",{"title":88,"url":89},"Braintrust pricing","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fog.jpg","ai-engineer",[94,95,96],"en","de","hu","DeepEval review: pytest for LLM outputs, and the judge bill","DeepEval runs LLM checks as pytest-style tests, with built-in judge metrics. The library is free and Apache-2.0; the judge calls and the data flow are the real cost.","Cover art for the DeepEval review: a test case passed through a judge metric to a pass or fail gate","Apache-2.0 · free; Confident AI from $0, Starter $200 a month","Evaluation framework","DeepEval review: pytest for LLM outputs · Balázs Csorba",[104,105,106,107,108],"DeepEval is an Apache-2.0 Python framework that runs LLM checks as unit tests, and the library works with no account.","Most predefined metrics use another model as the judge, so every run can cost judge tokens and can return a different score.","With default settings the judge is OpenAI, so test cases leave your network. A local judge such as Ollama keeps them inside, at a quality you have to measure.","Confident AI's cloud has a free plan, Starter at $200 a month and Team at $2,000 a month, with a DPA for every customer and an EU region on every plan.","Choose Ragas, Promptfoo or Braintrust for an Apache-2.0 metrics toolkit, a declarative CLI with red teaming, or a hosted UI with its own meter.",[110,113,116,119],{"q":111,"a":112},"Does DeepEval send my data anywhere?","Yes, to the judge model you configure. With no model set, the FAQ says the judge defaults to OpenAI. Results reach Confident AI only when you set a key, and the default storage region is the United States, with the EU available on every plan.",{"q":114,"a":115},"How much does DeepEval cost?","The library is free under an Apache-2.0 licence. Confident AI has a free plan with two seats and five test runs a week, Starter at $200 a month and Team at $2,000 a month, billed monthly per organisation. Judge calls are extra and depend on the model you choose.",{"q":117,"a":118},"Are the scores repeatable?","Not exactly. The G-Eval docs say the metric is not deterministic, so give thresholds some margin and re-run a failing case before you treat it as a regression. Tool Correctness is different, because its core score is deterministic.",{"q":120,"a":121},"Can I use a local model as the judge?","Yes. The metrics page lists Ollama among the judges, and you can wrap any other model by subclassing DeepEvalBaseLLM. Check the judge against cases you have labelled yourself, because a weaker judge moves every score.",[123,126,129,132,135,138,141,144,147],{"id":124,"title":125},"what-it-is","What it is",{"id":127,"title":128},"how-it-works","How it works",{"id":130,"title":131},"getting-started","Getting started",{"id":133,"title":134},"metrics-and-judge-cost","Metrics, and what the judge costs",{"id":136,"title":137},"goldens-red-teaming-and-ci","Goldens, red teaming and CI",{"id":139,"title":140},"cost-and-deployment","Cost and deployment",{"id":142,"title":143},"where-it-falls-short","Where it falls short",{"id":145,"title":146},"verdict","Verdict",{"id":148,"title":149},"sources","Sources",[151,155,179,182,189,209,210,213,222,225,226,233,235,236,239,292,299,300,303,310,333,335,336,339,385,388,398,405,407,410,416,417,420,421,424,442,443],{"type":152,"content":153},"paragraph",[154],"DeepEval is an open-source Python framework that runs LLM checks as unit tests, with more than thirty built-in metrics and an optional cloud platform called Confident AI. The verdict up front: take it if your engineers write Python, already run tests in CI and want model behaviour checked on every pull request. Skip it if the default OpenAI judge is not acceptable for your data, or if non-engineers need a hosted workspace before anything else.",{"type":152,"content":156},[157,158,163,164,168,169,173,174,178],"It sits next to ",{"tag":159,"to":160,"children":161},"link","\u002Ftools\u002Fragas",[162],"Ragas",", ",{"tag":159,"to":165,"children":166},"\u002Ftools\u002Fpromptfoo",[167],"Promptfoo"," and ",{"tag":159,"to":170,"children":171},"\u002Ftools\u002Fbraintrust",[172],"Braintrust",", which I have reviewed separately. For the process around the tests, from reading traces by hand to a gate in CI, see ",{"tag":159,"to":175,"children":176},"\u002Fblog\u002Fllm-evals-for-product-features",[177],"LLM evals for product features",".",{"type":180,"level":181,"id":124,"text":125},"heading",2,{"type":152,"content":183},[184,185,178],"DeepEval is the Apache-2.0 library behind Confident AI, which describes itself as the enterprise AI evals and observability platform. The library runs on your machine and needs no account. The current release on PyPI is 4.2.8, published on 2 October 2026, and it needs Python 3.9 or newer. Install it with ",{"tag":186,"children":187},"code",[188],"pip install -U deepeval",{"type":190,"ordered":191,"items":192},"list",false,[193,199,204],[194,198],{"tag":195,"children":196},"strong",[197],"Metrics. ","More than thirty named metrics in nine groups, from G-Eval and agent metrics to retrieval, multi-turn, safety and image checks.",[200,203],{"tag":195,"children":201},[202],"Judges. ","Almost all predefined metrics use an LLM as the judge, and you can point them at OpenAI, Azure OpenAI, Anthropic, Gemini, Ollama or LiteLLM.",[205,208],{"tag":195,"children":206},[207],"Red teaming. ","A separate Apache-2.0 package, DeepTeam, covers attacks and vulnerabilities.",{"type":180,"level":181,"id":127,"text":128},{"type":152,"content":211},[212],"A test is ordinary Python. Each metric receives a test case, which holds the input, the actual output and, depending on the metric, the expected output, the retrieval context or the tools the agent called. For a judge-based metric, the test case and the criteria go to the judge model, which returns a score from 0 to 1 and a reason. The threshold turns the score into a pass or a fail, and assert_test fails the test when the score is below it.",{"type":214,"attrs":215,"inner":219,"caption":220},"diagram",{"viewBox":216,"role":217,"aria-labelledby":218},"0 0 720 330","img","d1-de-t d1-de-d","\u003Ctitle id=\"d1-de-t\">How a DeepEval test decides pass or fail\u003C\u002Ftitle>\u003Cdesc id=\"d1-de-d\">A test case built from an input and an output goes to a metric. For judge-based metrics, the judge model returns a score from 0 to 1 with a reason. The threshold turns that score into a pass or a fail, pytest reports the result, and CI blocks the merge on a fail. Results reach Confident AI only when a key is set, and the optional upload is the dashed step.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-de\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">How a DeepEval test decides pass or fail\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">the judge call is the cost\u003C\u002Ftext>\u003Crect x=\"20\" y=\"52\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"95\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Test case\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"102\" text-anchor=\"middle\" class=\"d-small\">input and output\u003C\u002Ftext>\u003Crect x=\"195\" y=\"52\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"270\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Metric\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"102\" text-anchor=\"middle\" class=\"d-small\">G-Eval, built-in\u003C\u002Ftext>\u003Crect x=\"370\" y=\"52\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"445\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Judge model\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"102\" text-anchor=\"middle\" class=\"d-small\">default: OpenAI\u003C\u002Ftext>\u003Crect x=\"545\" y=\"52\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"620\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Score\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"102\" text-anchor=\"middle\" class=\"d-small\">0 to 1, with reason\u003C\u002Ftext>\u003Cpath d=\"M170 83 H193\" class=\"d-line\" marker-end=\"url(#ah-de)\" \u002F>\u003Cpath d=\"M345 83 H368\" class=\"d-line\" marker-end=\"url(#ah-de)\" \u002F>\u003Cpath d=\"M520 83 H543\" class=\"d-line\" marker-end=\"url(#ah-de)\" \u002F>\u003Cpath d=\"M620 114 V178\" class=\"d-line\" marker-end=\"url(#ah-de)\" \u002F>\u003Crect x=\"545\" y=\"180\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"620\" y=\"208\" text-anchor=\"middle\" class=\"d-text\">Threshold\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">pass at or above\u003C\u002Ftext>\u003Cpath d=\"M543 211 H522\" class=\"d-line\" marker-end=\"url(#ah-de)\" \u002F>\u003Crect x=\"370\" y=\"180\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"445\" y=\"208\" text-anchor=\"middle\" class=\"d-text\">pytest result\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">CI blocks a fail\u003C\u002Ftext>\u003Cpath d=\"M368 211 H347\" class=\"d-line d-dash\" marker-end=\"url(#ah-de)\" \u002F>\u003Crect x=\"195\" y=\"180\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"270\" y=\"208\" text-anchor=\"middle\" class=\"d-text\">Confident AI\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">optional upload\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"278\" class=\"d-small\">Without a key, every result stays on your machine.\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"302\" class=\"d-label\">The judge is the only step that calls a model for judge-based metrics.\u003C\u002Ftext>",[221],"For judge-based metrics, the judge model is the only step that calls a model, and results reach Confident AI only when a key is set.",{"type":152,"content":223},[224],"The judge is the part that matters for cost and reliability. The metrics page says almost all predefined metrics use an LLM as the judge, and that any judge can be used, including OpenAI, Azure OpenAI, Ollama, Anthropic, Gemini and LiteLLM. You can also wrap your own model by subclassing DeepEvalBaseLLM. The FAQ says the judge defaults to OpenAI when you specify no model.",{"type":180,"level":181,"id":130,"text":131},{"type":152,"content":227},[228,229,232],"Export an OpenAI key for the judge, install the package and write the test in a file such as test_example.py. The example checks an answer with G-Eval, where you describe the criteria in plain language and DeepEval writes the evaluation steps from them. Run it with ",{"tag":186,"children":230},[231],"deepeval test run test_example.py",". The test fails if the correctness score is below 0.5.",{"type":186,"code":234},"# test_example.py\nfrom deepeval import assert_test\nfrom deepeval.metrics import GEval\nfrom deepeval.test_case import LLMTestCase, SingleTurnParams\n\n\ndef test_refund_answer():\n    correctness = GEval(\n        name='Correctness',\n        criteria='Is the actual output a correct answer to the input, consistent with the expected output?',\n        evaluation_params=[\n            SingleTurnParams.INPUT,\n            SingleTurnParams.ACTUAL_OUTPUT,\n            SingleTurnParams.EXPECTED_OUTPUT,\n        ],\n        threshold=0.5,\n    )\n    test_case = LLMTestCase(\n        input='Can I return shoes after 30 days?',\n        actual_output='You can return them within 30 days of delivery, if they are unworn.',\n        expected_output='Returns are accepted within 30 days of delivery if the item is unworn.',\n    )\n    assert_test(test_case, [correctness])",{"type":180,"level":181,"id":133,"text":134},{"type":152,"content":237},[238],"Metric choice decides the bill. Faithfulness extracts the claims in an answer, checks each claim against the retrieval context, and scores the share of claims that do not contradict it, with a default threshold of 0.5. The judge has more to read when an answer makes many claims, so the cost grows with the answer. The evaluate() function brings caching, parallelisation, cost tracking and error handling, while a standalone measure() call loses those optimisations. The docs also warn that many metric calls at once can trigger rate-limit errors, so cap the concurrency to your provider's limit.",{"type":240,"head":241,"rows":246},"table",[242,244],[243],"Group",[245],"Named metrics",[247,252,257,262,267,272,277,282,287],[248,250],[249],"Custom",[251],"G-Eval, DAG, Arena G-Eval, JevEval, and your own code metrics such as BLEU or ROUGE",[253,255],[254],"Agents, trajectory",[256],"Task Completion, Step Efficiency, Plan Adherence, Plan Quality",[258,260],[259],"Agents, components",[261],"Tool Correctness, Argument Correctness",[263,265],[264],"Retriever",[266],"Contextual Relevancy, Contextual Precision, Contextual Recall",[268,270],[269],"Generator",[271],"Answer Relevancy, Faithfulness",[273,275],[274],"Multi-turn chatbots",[276],"Knowledge Retention, Role Adherence, Conversation Completeness, Conversation Relevancy",[278,280],[279],"Safety",[281],"Bias, Toxicity, Non-Advice, Misuse, PIILeakage, Role Violation",[283,285],[284],"Image",[286],"Image Coherence, Image Helpfulness, Image Reference, Text-to-Image, Image-Editing",[288,290],[289],"Other",[291],"Hallucination, JSON Correctness, Summarization, Ragas",{"type":293,"variant":294,"title":295,"body":296},"callout","note","Deterministic where it can be",[297],[298],"Tool Correctness is the exception to the judge-first rule. Its core score counts how many of the tools the agent called match the tools it was expected to call, and no model is involved. An optional second check uses an LLM to judge tool selection, but only when you pass the available tools, and the final score is the lower of the two.",{"type":180,"level":181,"id":136,"text":137},{"type":152,"content":301},[302],"Hand-written test cases run out quickly, so DeepEval can generate them. generate_goldens_from_docs takes a list of document paths, reads .txt, .docx, .pdf and Markdown files, stores the chunks in chromadb and has a critic model score each chunk from 0 to 1. By default each golden also gets an expected output. Without an OpenAI key you must supply your own embedding model and LLM, because the default embedder is text-embedding-3-small. The docs describe no review step, so have someone read a sample before a generated set becomes a regression suite.",{"type":152,"content":304},[305,306,309],"Red teaming is a separate package. DeepTeam is an Apache-2.0 framework to red team LLMs and AI agents, with more than 50 ready-made vulnerabilities and more than 20 research-backed attack methods, in single-turn and multi-turn form. Its LLM-as-a-judge metrics run on your machine and return a pass or fail with reasoning. It maps its checks to the OWASP Top 10 for LLMs 2025, the OWASP Top 10 for Agents 2026, NIST AI RMF and MITRE ATLAS, and it installs with ",{"tag":186,"children":307},[308],"pip install -U deepteam",". The platform's AI red-teaming module sits in the Enterprise ++ tier of the pricing page.",{"type":152,"content":311},[312,313,316,317,320,321,324,325,328,329,332],"The CI pattern in the docs is short. Store the judge key as ",{"tag":186,"children":314},[315],"OPENAI_API_KEY",", add ",{"tag":186,"children":318},[319],"CONFIDENT_API_KEY"," only if the results should reach the platform, and run ",{"tag":186,"children":322},[323],"deepeval test run",". Without the key the same tests still run locally. Add ",{"tag":186,"children":326},[327],"--official"," (or ",{"tag":186,"children":330},[331],"-o",") to mark a run as the official baseline on Confident AI.",{"type":186,"code":334},"- name: Run LLM tests\n  env:\n    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n    CONFIDENT_API_KEY: ${{ secrets.CONFIDENT_API_KEY }}\n  run: poetry run deepeval test run test_llm_app.py",{"type":180,"level":181,"id":139,"text":140},{"type":152,"content":337},[338],"The library costs nothing to run beyond the judge. The platform has four plans, billed monthly per organisation, and the pricing page carries no 'as of' date, so check it again before you budget.",{"type":240,"head":340,"rows":349},[341,343,345,347],[342],"Plan",[344],"Price",[346],"Included",[348],"What it adds",[350,359,368,377],[351,353,355,357],[352],"Free",[354],"$0",[356],"2 seats, 1 project, 1 GB-month of trace data",[358],"5 test runs a week",[360,362,364,366],[361],"Starter",[363],"$200 a month",[365],"Unlimited seats, 5 projects, 5 GB-months, then $1 per GB-month",[367],"Online evals, annotation queues, real-time alerts",[369,371,373,375],[370],"Team",[372],"$2,000 a month",[374],"Unlimited seats and projects, 75 GB-months, then $1 per GB-month",[376],"SOC 2, SSO, custom roles, custom contracts and SLAs",[378,380,381,383],[379],"Enterprise",[249],[382],"Unlimited usage",[384],"On-prem, custom data residency, 24x7 support, red-teaming module (Enterprise ++)",{"type":152,"content":386},[387],"For comparison, Braintrust's Starter tier is free with 14 days of retention, and its Pro plan is $249 a month, with processed data at $3 per GB and scores at $1.50 per 1,000 beyond the included amounts. The Confident AI Starter plan is a flat $200 a month with no per-seat fee, and it includes online evals and annotation queues.",{"type":152,"content":389},[390,391,393,394,397],"There are three separate data flows, and they answer different questions. The judge receives the test case, so with no model configured that content goes to OpenAI, from your laptop or CI runner. Confident AI receives results only when ",{"tag":186,"children":392},[319]," is set, and the GitHub README says that when you use the cloud platform all test cases are logged automatically. Telemetry is the third flow: by default DeepEval records basic, non-identifying counts, such as how many evaluations ran and which metrics were used, and the data-privacy page names PostHog as the only destination. Set ",{"tag":186,"children":395},[396],"DEEPEVAL_TELEMETRY_OPT_OUT=1"," to switch it off.",{"type":152,"content":399},[400,401,404],"On the platform, the FAQ says data is stored in a private AWS cloud that only your organisation can access. The default region is the United States, and the EU is available on every plan. Choose it at login or sign-up, or run ",{"tag":186,"children":402},[403],"deepeval set-confident-region EU",", which configures DeepEval only. The region page adds that an API key alone does not decide where the data goes, so set the two endpoints as well.",{"type":186,"code":406},"deepeval set-confident-region EU\nexport CONFIDENT_BASE_URL=https:\u002F\u002Feu.api.confident-ai.com\nexport CONFIDENT_OTEL_ENDPOINT=https:\u002F\u002Feu.otel.confident-ai.com\u002Fv1\u002Ftraces",{"type":152,"content":408},[409],"For the contract, the subprocessor list, last modified on 18 February 2026, says every customer is given a data processing agreement. Personal data stays in the chosen region unless it has to move for performance or availability, or as agreed. The list names OpenAI for AI model inference, US only, and AWS, ClickHouse, Supabase and PostHog as US and Europe. I could not find a retention period for the cloud in any page I opened, so ask for it in writing before production data goes in. Self-hosting is an Enterprise option, and the on-premise setup points the same two variables at your own hosts.",{"type":293,"variant":411,"title":412,"body":413},"warn","What the EU option does not change",[414],[415],"The region setting covers where Confident AI stores and processes platform data. It does not move the judge. If the judge is OpenAI, the calls still go to OpenAI, wherever your platform data lives. For a strict EU setup, host the judge yourself, for example with Ollama, and measure its scores against cases you have labelled before you rely on them.",{"type":180,"level":181,"id":142,"text":143},{"type":152,"content":418},[419],"Three weaknesses matter most. First, scores are not repeatable. The G-Eval docs say the metric is not deterministic, so a score near the threshold will sometimes flip, and the judge caps the quality of everything you measure with it. Second, the surface is large. The metric list is long, the document route for generated goldens needs chromadb plus several LangChain packages, and the docs show a TypeScript command next to the Python one, although everything I checked for this review is Python. Third, the platform's price steps are coarse. The Free plan's five test runs a week will not carry a busy pipeline, and the jump from $200 to $2,000 a month is large for a team that only wants shared history.",{"type":180,"level":181,"id":145,"text":146},{"type":152,"content":422},[423],"DeepEval is the right default for a Python team that wants LLM behaviour tests in the same runner and the same pull request as the rest of its suite. The library and DeepTeam are both Apache-2.0, the metric list covers agents, retrieval, conversations and safety without writing your own judges, and the local mode works with no account. I would not adopt it where the test data cannot reach a judge you trust, or where non-engineers need a hosted workspace from day one. Pick an alternative in those cases.",{"type":190,"ordered":425,"items":426},true,[427,432,437],[428,431],{"tag":195,"children":429},[430],"Ragas if ","you want an Apache-2.0 toolkit with pre-built metrics and test data generation, and you do not need the Confident AI platform.",[433,436],{"tag":195,"children":434},[435],"Promptfoo if ","you want declarative configs that compare prompts and models from the command line in CI, with red teaming in the same tool. Its MIT repository now says Promptfoo is part of OpenAI, so weigh that against your vendor rules.",[438,441],{"tag":195,"children":439},[440],"Braintrust if ","you want a hosted UI and a click-through DPA on Pro, and you accept its meter, where usage above the included data and scores is billed on top of the $249 a month.",{"type":180,"level":181,"id":148,"text":149},{"type":190,"ordered":191,"items":444},[445,449,452,455,458,461,464,467,470,473,476,479,482,485,488,491,494,497],[446],{"tag":447,"href":38,"children":448},"a",[37],[450],{"tag":447,"href":41,"children":451},[40],[453],{"tag":447,"href":44,"children":454},[43],[456],{"tag":447,"href":47,"children":457},[46],[459],{"tag":447,"href":50,"children":460},[49],[462],{"tag":447,"href":53,"children":463},[52],[465],{"tag":447,"href":56,"children":466},[55],[468],{"tag":447,"href":59,"children":469},[58],[471],{"tag":447,"href":62,"children":472},[61],[474],{"tag":447,"href":65,"children":475},[64],[477],{"tag":447,"href":68,"children":478},[67],[480],{"tag":447,"href":71,"children":481},[70],[483],{"tag":447,"href":74,"children":484},[73],[486],{"tag":447,"href":77,"children":487},[76],[489],{"tag":447,"href":80,"children":490},[79],[492],{"tag":447,"href":83,"children":493},[82],[495],{"tag":447,"href":86,"children":496},[85],[498],{"tag":447,"href":89,"children":499},[88],[501,614,691,797],{"slug":502,"published":5,"minutes":6,"category":7,"tags":503,"keywords":509,"about":517,"sources":530,"cover":606,"og":607,"expertise":92,"locales":608,"lang":94,"title":609,"description":610,"coverAlt":611,"url":520,"pricing":612,"kind":613},"dspy",[504,505,506,507,508],"Prompt optimisation","LLM programs","MIPROv2","GEPA","Python",[502,510,511,512,513,514,515,516],"dspy tutorial","dspy optimizer","miprov2","gepa prompt optimization","dspy vs prompt engineering","automatic prompt optimization","dspy signatures",[518,521,524,527],{"name":519,"url":520},"DSPy","https:\u002F\u002Fdspy.ai",{"name":522,"url":523},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",{"name":525,"url":526},"Bayesian optimization","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBayesian_optimization",{"name":528,"url":529},"Fine-tuning (deep learning)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FFine-tuning_(deep_learning)",[531,534,537,540,543,546,549,552,555,558,561,564,567,570,573,576,579,582,585,588,591,594,597,600,603],{"title":532,"url":533},"DSPy home and cost example","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002F",{"title":535,"url":536},"DSPy installation","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Finstallation\u002F",{"title":538,"url":539},"DSPy: program, don't prompt","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fprogram-dont-prompt\u002F",{"title":541,"url":542},"DSPy signatures","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fdiving-deeper\u002Fsignatures-in-depth\u002F",{"title":544,"url":545},"DSPy metrics","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fmetrics\u002F",{"title":547,"url":548},"DSPy Example API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fprimitives\u002FExample\u002F",{"title":550,"url":551},"DSPy Predict API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FPredict\u002F",{"title":553,"url":554},"DSPy ChainOfThought API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FChainOfThought\u002F",{"title":556,"url":557},"DSPy ReAct API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FReAct\u002F",{"title":559,"url":560},"DSPy optimizers index","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002F",{"title":562,"url":563},"DSPy MIPROv2 API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FMIPROv2\u002F",{"title":565,"url":566},"MIPROv2 source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fmipro_optimizer_v2.py",{"title":568,"url":569},"DSPy BootstrapFewShot API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFewShot\u002F",{"title":571,"url":572},"DSPy BootstrapFinetune API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFinetune\u002F",{"title":574,"url":575},"DSPy GEPA guide","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fgepa-optimization\u002F",{"title":577,"url":578},"GEPA source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fgepa\u002Fgepa.py",{"title":580,"url":581},"DSPy saving programs","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fsaving\u002F",{"title":583,"url":584},"DSPy caching","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fcache\u002F",{"title":586,"url":587},"DSPy LM API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodels\u002FLM\u002F",{"title":589,"url":590},"DSPy FAQ","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ffaqs\u002F",{"title":592,"url":593},"DSPy GitHub repository","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy",{"title":595,"url":596},"DSPy 3.4.0 release","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Freleases\u002Ftag\u002F3.4.0",{"title":598,"url":599},"DSPy paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.03714",{"title":601,"url":602},"MIPRO paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.11695",{"title":604,"url":605},"GEPA paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2507.19457","\u002Fimages\u002Fblog\u002Fdspy\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdspy\u002Fog.jpg",[94,95,96],"DSPy review: compile your prompts against a metric, not by hand","DSPy compiles prompts from signatures, a metric and examples. What the optimisers cost in model calls, when they pay off, and when a hand-written prompt wins.","Cover art for the DSPy review: a signature and a metric feed an optimiser that compiles a saved program.","MIT · free, you pay the model API","Prompt optimisation framework",{"slug":615,"published":616,"minutes":6,"category":7,"tags":617,"keywords":623,"about":631,"sources":641,"cover":683,"og":684,"expertise":92,"locales":685,"lang":94,"title":686,"description":687,"coverAlt":688,"url":633,"pricing":689,"kind":690},"llama-cpp","2026-10-05",[618,619,620,621,622],"llama.cpp","GGUF","Quantisation","Local inference","llama-server",[618,624,625,626,627,628,629,630],"llama.cpp review","GGUF quantisation levels","llama-server OpenAI compatible","llama.cpp vs Ollama","run LLM locally GDPR","llama.cpp CUDA Metal Vulkan","Q4_K_M vs Q8_0",[632,634,635,638],{"name":618,"url":633},"https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp",{"name":33,"url":34},{"name":636,"url":637},"Quantization (signal processing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FQuantization_(signal_processing)",{"name":639,"url":640},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",[642,644,647,650,653,656,659,662,665,668,671,674,677,680],{"title":643,"url":633},"llama.cpp repository: goals, backends, licence",{"title":645,"url":646},"llama.cpp releases: builds b11538 to b11541","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Freleases",{"title":648,"url":649},"llama.cpp build documentation: CMake flags","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Fdocs\u002Fbuild.md",{"title":651,"url":652},"llama.cpp server README: endpoints, slots and grammars","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fserver\u002FREADME.md",{"title":654,"url":655},"llama.cpp quantisation README: bits, size and speed","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fquantize\u002FREADME.md",{"title":657,"url":658},"GGUF specification in the ggml repository","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fggml\u002Fblob\u002Fmaster\u002Fdocs\u002Fgguf.md",{"title":660,"url":661},"Performance of llama.cpp on Apple Silicon M-series","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F4167",{"title":663,"url":664},"Performance of llama.cpp on Nvidia CUDA","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F15013",{"title":666,"url":667},"Ollama README: supported backends and REST API","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":669,"url":670},"LM Studio documentation: app overview","https:\u002F\u002Flmstudio.ai\u002Fdocs\u002Fapp",{"title":672,"url":673},"vLLM README: features, hardware and licence","https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm",{"title":675,"url":676},"vLLM documentation: GGUF support","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Flatest\u002Ffeatures\u002Fquantization\u002Fgguf.html",{"title":678,"url":679},"GDPR Article 28: processor","https:\u002F\u002Fgdpr-info.eu\u002Fart-28-gdpr\u002F",{"title":681,"url":682},"GDPR Article 32: security of processing","https:\u002F\u002Fgdpr-info.eu\u002Fart-32-gdpr\u002F","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fog.jpg",[94,95,96],"llama.cpp review: the local engine under Ollama and LM Studio","llama.cpp runs open models in plain C and C++ on Metal, CUDA, Vulkan or the CPU. I cover GGUF quants, llama-server and where it falls short.","Cover art for the llama.cpp review: one GGUF file fans out to Metal, CUDA, Vulkan and CPU, then one OpenAI-style API.","MIT · free","Local inference runtime",{"slug":692,"published":616,"minutes":693,"category":7,"tags":694,"keywords":700,"about":708,"sources":720,"cover":790,"og":791,"expertise":92,"locales":792,"lang":94,"title":793,"description":794,"coverAlt":795,"url":732,"pricing":796,"kind":695},"opik",7,[695,696,697,698,699],"LLM observability","Tracing","Evaluation","OpenTelemetry","Self-hosting",[692,701,702,703,704,705,706,707],"opik review","opik vs langfuse","opik pricing","self-hosted llm tracing","opik opentelemetry","llm as a judge metrics","opik guardrails",[709,712,714,717],{"name":710,"url":711},"Comet ML","https:\u002F\u002Fwww.comet.com",{"name":698,"url":713},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":715,"url":716},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",{"name":718,"url":719},"Apache License","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_License",[721,724,727,730,733,736,739,742,745,748,751,754,757,760,763,766,769,772,775,778,781,784,787],{"title":722,"url":723},"Opik repository README on GitHub","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik",{"title":725,"url":726},"Opik LICENSE file: Apache License 2.0","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik\u002Fblob\u002Fmain\u002FLICENSE",{"title":728,"url":729},"Opik on PyPI: package metadata","https:\u002F\u002Fpypi.org\u002Fpypi\u002Fopik\u002Fjson",{"title":731,"url":732},"Opik product page","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fproducts\u002Fopik\u002F",{"title":734,"url":735},"Opik pricing","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fpricing\u002F",{"title":737,"url":738},"Opik docs: run locally with Docker Compose","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Flocal_deployment",{"title":740,"url":741},"Opik docs: Kubernetes deployment with Helm","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Fkubernetes",{"title":743,"url":744},"Opik docs: platform architecture","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Farchitecture",{"title":746,"url":747},"Opik docs: tracing getting started","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Ftracing\u002Fgetting-started",{"title":749,"url":750},"Opik docs: OpenTelemetry integration","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fintegrations\u002Fopentelemetry",{"title":752,"url":753},"Opik docs: evaluation metrics overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fevaluation\u002Fmetrics\u002Foverview",{"title":755,"url":756},"Opik docs: online evaluation rules","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fonline-evaluation\u002Frules",{"title":758,"url":759},"Opik docs: Prompt Library overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Fprompt-library\u002Foverview",{"title":761,"url":762},"Opik docs: optimisation algorithms","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Foptimization-runs\u002Falgorithms\u002Foverview",{"title":764,"url":765},"Opik docs: guardrails overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Foverview",{"title":767,"url":768},"Opik docs: guardrails server","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Fserver",{"title":770,"url":771},"Opik docs: SDK anonymizers","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fgateway-guardrails\u002Fanonymizers",{"title":773,"url":774},"Opik docs: data anonymization","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fadministration\u002Fdata_anonymization",{"title":776,"url":777},"Comet ML privacy policy","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fprivacy-policy\u002F",{"title":779,"url":780},"Comet Trust Center","https:\u002F\u002Ftrust.comet.com\u002F",{"title":782,"url":783},"Langfuse repository README (licence)","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse",{"title":785,"url":786},"Arize Phoenix repository README (licence)","https:\u002F\u002Fgithub.com\u002FArize-ai\u002Fphoenix",{"title":788,"url":789},"LangChain pricing (LangSmith plans)","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fopik\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopik\u002Fog.jpg",[94,95,96],"Opik review: open-source tracing and evals, with a US-hosted cloud","Opik puts traces, LLM-as-a-judge metrics, prompt versions and an optimiser on one Apache 2.0 platform. Free to self-host, Pro cloud at $19 a month, US-hosted.","Cover art for the Opik review: a trace moves from the app through the backend to a judge and a score gate.","Apache 2.0 · free to self-host · Pro cloud from $19 a month",{"slug":798,"published":799,"minutes":800,"category":7,"tags":801,"keywords":804,"about":811,"sources":815,"cover":833,"og":834,"expertise":92,"locales":835,"lang":94,"title":836,"description":837,"coverAlt":838,"url":839,"pricing":840,"kind":690},"ollama","2026-09-29",11,[621,802,618,619,803],"Open models","Model serving",[798,805,806,807,808,809,810],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[812],{"name":813,"url":814},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[816,819,821,824,827,830],{"title":817,"url":818},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":820,"url":667},"Ollama on GitHub, with the MIT LICENSE file",{"title":822,"url":823},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":825,"url":826},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":828,"url":829},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":831,"url":832},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[94,95,96],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Cover art for the Ollama review: a request on port 11434 passes the scheduler and the engine and returns streamed tokens, with model loading and idle unloading noted below.","https:\u002F\u002Follama.com","MIT · free for personal use",1791636874663]