Tools/LLMOps & evals
Promptfoo: LLM evals and red teaming from one YAML file
Promptfoo in 2026: MIT-licensed evals and red teaming, now inside OpenAI. What it does well, where the YAML approach breaks, and what the tiers cost.
- Type
- Evaluation and red teaming
- Pricing
- MIT · Enterprise paid
Balázs Csorba··10 min read
- LLM evaluation
- Red teaming
- Prompt testing
- CI/CD

Key takeaways
- Promptfoo is a MIT-licensed CLI that runs quality evals and red teaming from one YAML file, and the repository declared itself part of OpenAI after the March 2026 acquisition announcement.
- The Community tier is free and includes 10,000 red team probes a month; dashboards, custom plugins, SSO and the API are quote-only Enterprise rows.
- Declarative test cases are the real strength, because a reviewer can read an eval change as a diff without running a Python test suite.
- Model-graded assertions such as llm-rubric are nondeterministic and cost a judge call per cell, so the pass rate measures the grader as much as the prompt.
- Keep the config and the test data portable: MIT keeps the code forkable, but a model-comparison tool owned by one candidate is a neutrality problem the licence does not solve.
Promptfoo is an open-source CLI and library that runs LLM evaluations and adversarial red teaming from a single YAML file. Position up front: it is the most practical way to turn a prompt change or a model swap into a text diff that CI can fail on, and since March 2026 it is also an eval tool that belongs to OpenAI.
It sits between a folder of pytest scripts and a hosted evaluation platform. Compared with LangSmith and Braintrust it gives up dashboards, shared history and a team UI in the free tier, and in exchange it runs on your machine, talks directly to more than 60 providers and needs no account. Compared with DeepEval it gives up Python expressiveness for a config file that a non-programmer can read in review.
What it is
Two products share one repository and one config format. The evaluation side runs prompts against test cases and grades the outputs; the red teaming side attacks a configured target with generated payloads and files the results as a vulnerability report. Both are driven from promptfooconfig.yaml, both run from the same command line and both can be wired into CI.
- MIT-licensed CLI and library: version 0.124.0 on npm, about 25.8k stars and 2.4k forks on GitHub
- Declarative configuration: prompts, providers, tests and assertions in one YAML file
- More than 60 providers, including OpenAI, Anthropic, Google, Azure, Bedrock, Ollama and custom HTTP, Python or JavaScript endpoints
- Two commands cover most of the work:
promptfoo evalfor quality,promptfoo redteam runfor security - The Community tier is free and includes 10,000 red team probes a month; Enterprise and On-Premise are priced on request
- The vendor reports 350k developers, 130k monthly actives and more than 25% of the Fortune 500
- The site footer carries SOC 2 and ISO 27001 badges, and the documentation ships a GitHub Action for CI
What the free tier deliberately withholds is the hosted part. There is no team dashboard, no searchable scan history, no continuous monitoring and no API in Community; those are Enterprise rows in the feature comparison. The tool itself stays a local process that calls your providers and writes results to your disk.
How it works
An evaluation is a matrix. The config declares prompts, providers and test cases, the runner expands the cross product, sends every cell to every model and scores the response against the assertions attached to that cell. Assertions are typed: substring and JSON schema checks, similarity, cost and latency thresholds, and model-graded checks such as llm-rubric that spend an extra model call to judge the output.
The runner is concurrent and caches completed calls, so a rerun after a small edit mostly re-reads the cache. Results open in a local web viewer as a side-by-side matrix, and the CLI can write JSON, YAML, CSV or HTML. The exit codes are the integration point: promptfoo eval returns 100 when at least one test fails or the pass rate falls below PROMPTFOO_PASS_RATE_THRESHOLD, and 1 for any other error.
Getting started
Install it with npm, brew or pip, or skip the install with npx promptfoo@latest. A minimal config that compares two models on one translated sentence:
# promptfooconfig.yaml: two models, one prompt, three checks
prompts:
- 'Translate this sentence into {{language}}: {{input}}'
providers:
- openai:gpt-6-sol
- openai:gpt-6-luna
tests:
- vars:
language: French
input: Hello world
assert:
- type: contains
value: 'Bonjour'
- type: llm-rubric
value: 'the translation is idiomatic and complete'
- type: cost
threshold: 0.01promptfoo eval runs every cell of that matrix and prints a pass rate; promptfoo view opens the side-by-side comparison in a browser. The -r flag swaps providers from the command line without editing the file, which is how the documentation suggests checking a candidate model against a baseline.
Red teaming
promptfoo redteam init writes a red team config, redteam setup walks through the target, redteam run generates and executes the attacks, and redteam report renders the findings. The documentation describes each plugin as a trained model that produces payloads for a specific weakness, rather than a fixed wordlist.
- 157 plugins across six categories: brand, compliance and legal, dataset, security and access control, trust and safety, and custom
- Security plugins map to the OWASP Top 10 for LLMs, the OWASP API Security Top 10 and MITRE ATLAS
- Compliance plugins map to NIST AI RMF, ISO/IEC 42001, GDPR Article 5 and EU AI Act Article 5
- Plugins are selected per target in the config, so a scan can be narrowed to the categories one application actually exposes
Framework mappings
For a team that has to show coverage rather than a pile of findings, the framework IDs are the useful part: nist:ai:measure:1.1, owasp:llm:01, iso:42001:privacy, gdpr:art5, eu:ai-act:art5. A report grouped by those IDs is a coverage argument a reviewer can follow. This is compliance mechanics rather than a compliance claim: the tool shows which probes ran, not that the product satisfies the regulation.
The free ceiling is 10,000 probes a month, and the plugins that need inference to generate and grade payloads are what consume it. Past that, the pricing page points at Enterprise for custom limits.
Who owns it now
On 9 March 2026 Promptfoo announced that it had agreed to be acquired by OpenAI. The post promised that the product would remain open source and MIT licensed and that the team would keep serving customers, and it noted that closing was subject to customary closing conditions. The repository now carries the line that Promptfoo is part of OpenAI, and the documentation was still being updated on 7 October 2026.
The licence keeps the code forkable; it does not keep the roadmap neutral. The engineering worry is narrow and concrete: a tool whose job is to answer which model should we ship is now owned by one of the candidates, and its red teaming, guardrails and model security products sit next to OpenAI's agent platform. Nothing in the MIT licence prevents a fork, and nothing in it obliges the company to prioritise multi-model comparison either.
Where it shingles
The first weakness is the config itself. A file with a dozen prompts and forty test cases is readable; one with four hundred cases is not, and anything that needs a loop, a join against production data or a custom parser pushes you into JavaScript assertion callbacks, at which point the language you avoided is back. The second is grading drift: model-graded assertions move between judge versions, so a green suite can turn amber because the grader moved. The third is scope, because a security team that needs an audit trail buys Enterprise regardless of how good the evals are.
| Tool | Interface | Free tier | Cost model |
|---|---|---|---|
| promptfoo | YAML and CLI, runs locally | Full evals, 10k red team probes a month | MIT; Enterprise quoted |
| LangSmith | Hosted platform plus SDK | 1 seat and 5k base traces a month | Plus from 39 USD per seat a month |
| Braintrust | Hosted platform plus SDK | Unlimited users, 1 GB processed data | Pro at 249 USD a month |
| DeepEval | Python, pytest style | The whole framework | Apache 2.0; hosted platform behind it |
The commercial shape is a free harness with a paid control plane. Continuous monitoring, a central dashboard, custom plugins, organisation-wide attack profiles, SSO, saved targets, searchable history and the API are Enterprise rows, and on-premise deployment adds a second quote. A solo developer never needs them; a team that has to prove what it tested last quarter does.
Verdict
Promptfoo is a good local harness that has grown a serious security half. It is fast, reviewable and honest about what it measures, and the red teaming coverage is now specific enough to argue with. The open question is not the licence, which stays MIT, but who decides which models the tool is tuned to notice.
- Take it if you want prompt and model changes gated in CI as text diffs, with an exit code that fails the job.
- Take it if you need adversarial testing before release and want it reviewed in the same pull request as your quality evals.
- Skip it if you need a hosted service, because dashboards, shared history, scan search and the API are Enterprise-only.
- Skip it if your evaluation logic needs real code, because a Python framework costs less effort than a YAML file fighting to express a loop.
- Keep your test data portable: the config is MIT and forkable, but a model-comparison tool owned by one of the candidates is a neutrality problem the licence does not solve.
A suite you can read in a pull request beats a dashboard you have to click through, and promptfoo is built on that premise. Its new owner is the reason to keep the exit code and the test data portable.
Sources
- Promptfoo documentation: intro
- Promptfoo pricing: plan comparison
- Promptfoo documentation: getting started
- Promptfoo documentation: command line usage (exit codes)
- Promptfoo documentation: red team plugins
- Promptfoo documentation: assertions and metrics
- Promptfoo repository on GitHub (MIT licence)
- Promptfoo blog: Promptfoo is joining OpenAI (9 March 2026)
Frequently asked questions
Is promptfoo free to use?
Yes. The Community version is MIT-licensed, runs locally and includes all evaluation features plus up to 10,000 red team probes per month. Enterprise adds custom probe limits, team sharing, continuous monitoring, SSO and API access, and both Enterprise and On-Premise are priced on request.
What is the difference between promptfoo evals and red teaming?
Evals run your prompts against test cases and grade the outputs with assertions such as substring checks, cost thresholds or llm-rubric. Red teaming generates adversarial payloads against a configured target: promptfoo redteam run executes the probes and promptfoo redteam report turns the findings into a vulnerability report.
Which model providers does promptfoo support?
The documentation lists more than 60 providers, including OpenAI, Anthropic, Google, Azure, Bedrock and local models through Ollama. Anything else can be reached through a custom HTTP, Python or JavaScript provider, so the config file does not have to change when the model behind an endpoint does.
What happened to promptfoo after the OpenAI acquisition?
Promptfoo announced on 9 March 2026 that it had agreed to be acquired by OpenAI, stating that the product would remain open source and MIT licensed and that the team would continue to serve customers. The repository now describes promptfoo as part of OpenAI, and the documentation is still being updated.