Tools/LLMOps & evals

Promptfoo: LLM evals and red teaming from one YAML file

Promptfoo in 2026: MIT-licensed evals and red teaming, now inside OpenAI. What it does well, where the YAML approach breaks, and what the tiers cost.

Type
Evaluation and red teaming
Pricing
MIT · Enterprise paid

··10 min read

  • LLM evaluation
  • Red teaming
  • Prompt testing
  • CI/CD
Diagram: a promptfooconfig.yaml file feeds a runner that calls every provider, assertions grade each output, and the results land in a matrix for the viewer and the CI gate.

Key takeaways

  • Promptfoo is a MIT-licensed CLI that runs quality evals and red teaming from one YAML file, and the repository declared itself part of OpenAI after the March 2026 acquisition announcement.
  • The Community tier is free and includes 10,000 red team probes a month; dashboards, custom plugins, SSO and the API are quote-only Enterprise rows.
  • Declarative test cases are the real strength, because a reviewer can read an eval change as a diff without running a Python test suite.
  • Model-graded assertions such as llm-rubric are nondeterministic and cost a judge call per cell, so the pass rate measures the grader as much as the prompt.
  • Keep the config and the test data portable: MIT keeps the code forkable, but a model-comparison tool owned by one candidate is a neutrality problem the licence does not solve.

Promptfoo is an open-source CLI and library that runs LLM evaluations and adversarial red teaming from a single YAML file. Position up front: it is the most practical way to turn a prompt change or a model swap into a text diff that CI can fail on, and since March 2026 it is also an eval tool that belongs to OpenAI.

It sits between a folder of pytest scripts and a hosted evaluation platform. Compared with LangSmith and Braintrust it gives up dashboards, shared history and a team UI in the free tier, and in exchange it runs on your machine, talks directly to more than 60 providers and needs no account. Compared with DeepEval it gives up Python expressiveness for a config file that a non-programmer can read in review.

What it is

Two products share one repository and one config format. The evaluation side runs prompts against test cases and grades the outputs; the red teaming side attacks a configured target with generated payloads and files the results as a vulnerability report. Both are driven from promptfooconfig.yaml, both run from the same command line and both can be wired into CI.

  • MIT-licensed CLI and library: version 0.124.0 on npm, about 25.8k stars and 2.4k forks on GitHub
  • Declarative configuration: prompts, providers, tests and assertions in one YAML file
  • More than 60 providers, including OpenAI, Anthropic, Google, Azure, Bedrock, Ollama and custom HTTP, Python or JavaScript endpoints
  • Two commands cover most of the work: promptfoo eval for quality, promptfoo redteam run for security
  • The Community tier is free and includes 10,000 red team probes a month; Enterprise and On-Premise are priced on request
  • The vendor reports 350k developers, 130k monthly actives and more than 25% of the Fortune 500
  • The site footer carries SOC 2 and ISO 27001 badges, and the documentation ships a GitHub Action for CI

What the free tier deliberately withholds is the hosted part. There is no team dashboard, no searchable scan history, no continuous monitoring and no API in Community; those are Enterprise rows in the feature comparison. The tool itself stays a local process that calls your providers and writes results to your disk.

How it works

An evaluation is a matrix. The config declares prompts, providers and test cases, the runner expands the cross product, sends every cell to every model and scores the response against the assertions attached to that cell. Assertions are typed: substring and JSON schema checks, similarity, cost and latency thresholds, and model-graded checks such as llm-rubric that spend an extra model call to judge the output.

How one promptfoo evaluation runsA four stage pipeline: the config file feeds a runner that calls every provider in parallel with caching, assertions grade each output, and the results go to a web viewer and a CI exit code. Below, a matrix of prompts against providers with pass and fail cells.One config, one runpromptfoo evalconfigprompts, testsrunnerparallel, cachedassertionsgrade each outputviewer, CImatrix, exit codeRESULTS MATRIXevery cell is one prompt x test x providergreen passes, gold fails, grey is cachedexit code gates the pull request100 on a failed test, 1 on any other error
The promptfoo evaluation matrix: every prompt runs against every provider, assertions grade each cell, and the exit code is what CI reacts to.

The runner is concurrent and caches completed calls, so a rerun after a small edit mostly re-reads the cache. Results open in a local web viewer as a side-by-side matrix, and the CLI can write JSON, YAML, CSV or HTML. The exit codes are the integration point: promptfoo eval returns 100 when at least one test fails or the pass rate falls below PROMPTFOO_PASS_RATE_THRESHOLD, and 1 for any other error.

Getting started

Install it with npm, brew or pip, or skip the install with npx promptfoo@latest. A minimal config that compares two models on one translated sentence:

# promptfooconfig.yaml: two models, one prompt, three checks
prompts:
  - 'Translate this sentence into {{language}}: {{input}}'

providers:
  - openai:gpt-6-sol
  - openai:gpt-6-luna

tests:
  - vars:
      language: French
      input: Hello world
    assert:
      - type: contains
        value: 'Bonjour'
      - type: llm-rubric
        value: 'the translation is idiomatic and complete'
      - type: cost
        threshold: 0.01

promptfoo eval runs every cell of that matrix and prints a pass rate; promptfoo view opens the side-by-side comparison in a browser. The -r flag swaps providers from the command line without editing the file, which is how the documentation suggests checking a candidate model against a baseline.

Red teaming

promptfoo redteam init writes a red team config, redteam setup walks through the target, redteam run generates and executes the attacks, and redteam report renders the findings. The documentation describes each plugin as a trained model that produces payloads for a specific weakness, rather than a fixed wordlist.

  • 157 plugins across six categories: brand, compliance and legal, dataset, security and access control, trust and safety, and custom
  • Security plugins map to the OWASP Top 10 for LLMs, the OWASP API Security Top 10 and MITRE ATLAS
  • Compliance plugins map to NIST AI RMF, ISO/IEC 42001, GDPR Article 5 and EU AI Act Article 5
  • Plugins are selected per target in the config, so a scan can be narrowed to the categories one application actually exposes

Framework mappings

For a team that has to show coverage rather than a pile of findings, the framework IDs are the useful part: nist:ai:measure:1.1, owasp:llm:01, iso:42001:privacy, gdpr:art5, eu:ai-act:art5. A report grouped by those IDs is a coverage argument a reviewer can follow. This is compliance mechanics rather than a compliance claim: the tool shows which probes ran, not that the product satisfies the regulation.

The free ceiling is 10,000 probes a month, and the plugins that need inference to generate and grade payloads are what consume it. Past that, the pricing page points at Enterprise for custom limits.

Who owns it now

On 9 March 2026 Promptfoo announced that it had agreed to be acquired by OpenAI. The post promised that the product would remain open source and MIT licensed and that the team would keep serving customers, and it noted that closing was subject to customary closing conditions. The repository now carries the line that Promptfoo is part of OpenAI, and the documentation was still being updated on 7 October 2026.

The licence keeps the code forkable; it does not keep the roadmap neutral. The engineering worry is narrow and concrete: a tool whose job is to answer which model should we ship is now owned by one of the candidates, and its red teaming, guardrails and model security products sit next to OpenAI's agent platform. Nothing in the MIT licence prevents a fork, and nothing in it obliges the company to prioritise multi-model comparison either.

Where it shingles

The first weakness is the config itself. A file with a dozen prompts and forty test cases is readable; one with four hundred cases is not, and anything that needs a loop, a join against production data or a custom parser pushes you into JavaScript assertion callbacks, at which point the language you avoided is back. The second is grading drift: model-graded assertions move between judge versions, so a green suite can turn amber because the grader moved. The third is scope, because a security team that needs an audit trail buys Enterprise regardless of how good the evals are.

ToolInterfaceFree tierCost model
promptfooYAML and CLI, runs locallyFull evals, 10k red team probes a monthMIT; Enterprise quoted
LangSmithHosted platform plus SDK1 seat and 5k base traces a monthPlus from 39 USD per seat a month
BraintrustHosted platform plus SDKUnlimited users, 1 GB processed dataPro at 249 USD a month
DeepEvalPython, pytest styleThe whole frameworkApache 2.0; hosted platform behind it

The commercial shape is a free harness with a paid control plane. Continuous monitoring, a central dashboard, custom plugins, organisation-wide attack profiles, SSO, saved targets, searchable history and the API are Enterprise rows, and on-premise deployment adds a second quote. A solo developer never needs them; a team that has to prove what it tested last quarter does.

Verdict

Promptfoo is a good local harness that has grown a serious security half. It is fast, reviewable and honest about what it measures, and the red teaming coverage is now specific enough to argue with. The open question is not the licence, which stays MIT, but who decides which models the tool is tuned to notice.

  1. Take it if you want prompt and model changes gated in CI as text diffs, with an exit code that fails the job.
  2. Take it if you need adversarial testing before release and want it reviewed in the same pull request as your quality evals.
  3. Skip it if you need a hosted service, because dashboards, shared history, scan search and the API are Enterprise-only.
  4. Skip it if your evaluation logic needs real code, because a Python framework costs less effort than a YAML file fighting to express a loop.
  5. Keep your test data portable: the config is MIT and forkable, but a model-comparison tool owned by one of the candidates is a neutrality problem the licence does not solve.
A suite you can read in a pull request beats a dashboard you have to click through, and promptfoo is built on that premise. Its new owner is the reason to keep the exit code and the test data portable.

Sources

  1. Promptfoo documentation: intro
  2. Promptfoo pricing: plan comparison
  3. Promptfoo documentation: getting started
  4. Promptfoo documentation: command line usage (exit codes)
  5. Promptfoo documentation: red team plugins
  6. Promptfoo documentation: assertions and metrics
  7. Promptfoo repository on GitHub (MIT licence)
  8. Promptfoo blog: Promptfoo is joining OpenAI (9 March 2026)

Frequently asked questions

Is promptfoo free to use?

Yes. The Community version is MIT-licensed, runs locally and includes all evaluation features plus up to 10,000 red team probes per month. Enterprise adds custom probe limits, team sharing, continuous monitoring, SSO and API access, and both Enterprise and On-Premise are priced on request.

What is the difference between promptfoo evals and red teaming?

Evals run your prompts against test cases and grade the outputs with assertions such as substring checks, cost thresholds or llm-rubric. Red teaming generates adversarial payloads against a configured target: promptfoo redteam run executes the probes and promptfoo redteam report turns the findings into a vulnerability report.

Which model providers does promptfoo support?

The documentation lists more than 60 providers, including OpenAI, Anthropic, Google, Azure, Bedrock and local models through Ollama. Anything else can be reached through a custom HTTP, Python or JavaScript provider, so the config file does not have to change when the model behind an endpoint does.

What happened to promptfoo after the OpenAI acquisition?

Promptfoo announced on 9 March 2026 that it had agreed to be acquired by OpenAI, stating that the product would remain open source and MIT licensed and that the team would continue to serve customers. The repository now describes promptfoo as part of OpenAI, and the documentation is still being updated.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.