Tools/LLMOps & evals

Braintrust review: eval-first observability with a hard meter

Braintrust turns production traces into datasets and gated experiments. What Starter and Pro really include, which parts are open source, and where Phoenix, Langfuse and LangSmith win.

Type
Evaluation platform
Pricing
Free · from $249 per month

··10 min read

  • Evaluations
  • Observability
  • LLM-as-a-judge
  • CI gates
  • Tracing
A loop from instrumented application logs into a dataset, an experiment with scorers, and a comparison that gates the pull request before the change returns to the application.

Key takeaways

  • Starter is genuinely usable but small: 1 GB of processed data, 10,000 scores and 14 days of retention, then Pro at $249 a month for 5 GB, 50,000 scores and 30 days.
  • Processed data is measured at ingestion, so deleting traces does not reduce the bill, and overage runs $4 per GB on Starter and $3 per GB on Pro.
  • The platform is closed source while the tooling is open: SDKs for six languages under Apache-2.0, autoevals under MIT and the bt CLI under Apache-2.0.
  • BYOC and self-hosted deployment require Enterprise, and SAML SSO, audit logging, custom retention and SOC 2 attestation are Enterprise rows as well.
  • There is no tier between $0 and $249, which makes Braintrust cheap for a solo evaluation project and expensive for a team that only wanted dashboards.

Braintrust is a hosted platform for instrumenting an LLM application, keeping its traces and scoring them against datasets in a way that repeats. Its centre of gravity is the evaluation: an Eval() call that runs a task over a dataset, scores the outputs and stores every run as a comparable experiment. The position of this review is that Braintrust is the strongest eval-first product in its class and the most tightly metered of its peers — the engineering is excellent, and the billing rewards teams who know exactly how much data they intend to send.

It covers three jobs that are usually split between tools: tracing in production, offline evaluation, and the gate between them in continuous integration. Arize Phoenix and Langfuse compete on the open-source side, LangSmith for shops already building on LangGraph. What it tends to replace is worse: a set of test cases that call the model, print scores to stdout and get deleted the week they turn flaky.

What it is

Braintrust is a SaaS product from Braintrust Data, built on three objects: logs, which are traces from the application; datasets, which are rows of input and expected output; and experiments, which run a task over a dataset with scorers attached. The documentation describes a five-step workflow — instrument, observe, annotate, evaluate, deploy — and ships SDKs for Python, TypeScript, Go, Ruby, Java and C#. The platform itself is closed source; much of the tooling around it is not.

  • Licence: platform closed source; SDKs Apache-2.0, autoevals MIT, braintrust-proxy MIT
  • Instrumentation: braintrust.auto_instrument() wraps the provider clients already in the process, plus a bt CLI for setup from a coding agent
  • Evaluation: Eval() over a dataset with autoevals scorers, custom code or LLM-as-a-judge
  • Plans: Starter at $0, Pro at $249 a month and Enterprise by quotation, all with unlimited users, projects, datasets and experiments
  • Metering: processed data in gigabytes, scores in thousands, model credits in dollars
  • Deployment: SaaS on Starter and Pro, with BYOC and self-hosted deployment only on Enterprise
  • Extras: playgrounds, environments and custom charts on Pro and above, and Loop, a built-in agent that writes scorers, test cases and prompt iterations

That last row is the tell. Braintrust assumes the platform is someone else's problem and the evaluation is yours, and teams that accept the premise get a very short loop from a production failure to a dataset row to a gated release. Teams that need the whole system inside their own account find that the price of that loop is an Enterprise contract, not the $249 tier.

How it works

Instrumentation happens in the process: braintrust.auto_instrument() patches the provider client the application already uses, so spans arrive with token costs attached and with no sidecar to deploy. Everything lands in a project, and logs, datasets and experiments share that project, which is what makes a production trace promotable into a test case.

The Braintrust loop from trace to gated releaseAn instrumented application logs traces and their costs to a project. Production traces become rows of a dataset, an experiment runs a task over that dataset with scorers attached, and a comparison of two runs gates the pull request before the change goes back to the application.One loop, three artefactsproject scopeApp + SDKauto_instrumentLogstraces, costDatasetrows from tracesExperimenttask + scorersCompare runsgate the pull requestScorers may be autoevals, custom code or a judge call that spends a model request per row.Every experiment is a permanent record: the second run is diffed against the first
Traces, datasets and experiments share one project, so a production failure can become a test case without an export.

The distinctive part is comparison. An experiment is a permanent record, and a second run under a new experiment name is diffed against the first in the interface. The documentation's own quickstart leans on this: the first run scores near zero under ExactMatch because the model answers with a sentence, the prompt is tightened to return a bare title, and the second experiment reports the gain as a score delta rather than an impression.

Getting started

Two keys and a file: BRAINTRUST_API_KEY for the platform, your provider key for the model, and one Python file holding data, task and scorer. The snippet below is the documented quickstart shape with the repetition trimmed out:

# export BRAINTRUST_API_KEY=...  and  export OPENAI_API_KEY=...
# pip install braintrust openai autoevals
import braintrust
from braintrust import Eval
from autoevals import ExactMatch
from openai import OpenAI

braintrust.auto_instrument()  # traces every call this process makes
client = OpenAI()

def task(input):
    r = client.responses.create(
        model="gpt-5-mini",
        input=[{"role": "system", "content": "Name the film."},
               {"role": "user", "content": input}],
    )
    return r.output_text

Eval(
    "movie-matcher",
    experiment_name="baseline-v1",
    data=[{"input": "A detective hunts a killer through the seven deadly sins.",
           "expected": "Se7en"}],
    task=task,
    scores=[ExactMatch],
)

Run it with bt eval movie_matcher.py and the terminal prints a link to the experiment; plain Python works as well, because the CLI only wraps the entry point. That same file is what a pull request runs: the eval-action GitHub workflow executes it against main and fails the build when a score regresses.

What it costs

Three plans and no middle step: the pricing page lists Starter at $0, Pro at $249 a month and Enterprise by quotation, with users, projects, datasets, playgrounds and experiments unlimited on all three. What differs is the meter:

PlanPriceProcessed dataScoresRetention
Starter$01 GB, then $4/GB10,000, then $2.50 per 1,00014 days
Pro$249 a month5 GB, then $3/GB50,000, then $1.50 per 1,00030 days, then $0.50 per GB-month
EnterpriseCustomCustomCustomCustom, up to 365 days

Model credits run alongside: $10 a month on Starter and $100 on Pro cover built-in models and platform features such as Topics, after which token rates apply. The tier jump is binary — a team that outgrows 1 GB and 14 days pays the full $249, though the docs offer six to twelve months of Pro to qualifying new startup customers.

What is open and what is not

The split deserves its own section, because the sentence Braintrust is open source is wrong in the way that matters: below Enterprise you cannot run the platform yourself. What is open is the instrumentation and the scoring tooling around it.

  • braintrust-sdk-python, -javascript, -go and -ruby: Apache-2.0
  • autoevals: MIT, about a thousand stars, the scorer library the quickstart imports
  • braintrust-proxy: MIT, a proxy in front of model traffic
  • agentbehavior: Apache-2.0, an open standard and dataset for agent behaviour
  • The bt CLI and the coding-agent plugins: Apache-2.0 or MIT; the hosted web application is not published

For an engineering team that distinction is mostly philosophical until procurement asks, and then it is the whole discussion. It also sets the exit path: the SDKs can be repointed at another backend without rewriting the application, but datasets, experiments and dashboards live in Braintrust's service, and scheduled export to S3 or Google Cloud Storage is an Enterprise feature.

Where it falls short

The first weakness is arithmetic. There is no tier between $0 and $249, the free tier's 1 GB and 14 days are a demo budget for a production agent, and scoring meters on top of ingestion, so a team running a judge across every production trace hits the score quota before the data quota. The second is what the lower tiers do not carry: SAML single sign-on, audit logging, custom retention, export automations, SOC 2 attestation and a business associate agreement are Enterprise rows, while Starter is limited to the owner permission group and one human-review score per project. The third is contractual drift: the documentation still carries a legacy-plan note for accounts created before 16 March 2026, so an evaluation done under the old terms is worth redoing.

ToolFree tierPaid entrySelf-host
Braintrust1 GB, 10,000 scores, 14 daysPro $249 a monthEnterprise only
Arize PhoenixNo caps on a local installAX Pro $50 a monthFree under ELv2
Langfuse50,000 units, 30 days, 2 usersCore $29 a monthFree, Docker Compose
LangSmith1 seat, 5,000 base tracesPlus $39 a seatEnterprise add-on

Read honestly, Braintrust and Phoenix are not competing on the same axis: one is priced for teams buying an evaluation process, the other for teams owning the infrastructure. Langfuse is the closest direct substitute at roughly a third of the entry price, and LangSmith is the natural pick when the application already sits on LangGraph and first-party trace formats matter more than the scorer library.

SaaS is available on all plans. BYOC and self-hosted deployments, which keep data in your own cloud, require Enterprise. — Braintrust documentation, plans and limits

Verdict

Braintrust earns its price when evaluation is a recurring engineering ritual rather than a one-off script: the experiment model, the scorer library and the CI action form a loop that is tedious to assemble from parts. It is a poor fit for a team whose main need is trace storage behind a dashboard, because that is a $249 habit with a data bill attached to it.

  1. Choose it when a pull request should fail on an eval regression and you want the scorer library to already exist.
  2. Choose it when several roles — engineers, reviewers, a product manager — need the same experiments and dashboards; seats are unlimited on every plan.
  3. Choose it when volume is predictable: 5 GB and 50,000 scores a month is a known quantity, and the overage rates are published.
  4. Do not choose it when traces must stay inside your own account; BYOC and self-hosting need an Enterprise contract.
  5. Do not choose it when the workload is chatty and unbounded: metering at ingestion plus per-score charges makes a noisy agent expensive in a way per-seat tools are not.

Sources

  1. Braintrust pricing: plans, credits and usage rates
  2. Braintrust documentation: plans and limits
  3. Braintrust documentation: evaluation quickstart
  4. Braintrust documentation: get started
  5. GitHub: Braintrust organisation repositories
  6. Arize Phoenix documentation: self-hosting
  7. Langfuse pricing: cloud plans and billable units
  8. LangChain pricing: LangSmith plans

Frequently asked questions

How much does Braintrust cost?

Starter is $0 with $10 of model credits, 1 GB of processed data, 10,000 scores and 14 days of retention, plus unlimited users, projects, datasets, playgrounds and experiments. Pro is $249 a month for $100 of credits, 5 GB, 50,000 scores, 30 days of retention, custom charts, environments, RBAC and priority support. Enterprise is custom-priced and adds BYOC, self-hosted deployment, SAML or OIDC single sign-on, audit logging, a BAA and an uptime SLA.

Is Braintrust open source?

The platform is not; much of the tooling is. The Python, TypeScript, Go, Ruby, Java and C# SDKs are Apache-2.0, autoevals is MIT with about a thousand GitHub stars, braintrust-proxy is MIT and the agentbehavior standard is Apache-2.0. There is no self-hosted edition of the service below the Enterprise plan.

What counts as a score in the pricing?

A score is one recorded evaluation result, whether it came from an LLM-as-a-judge call, an autoevals scorer or custom code, attached to a trace or an experiment row. Starter includes 10,000 a month and then charges $2.50 per thousand; Pro includes 50,000 and then $1.50 per thousand. Judge-based scoring multiplies quickly across a dataset, so the score quota is often reached before the data quota.

Braintrust or Arize Phoenix?

Choose Braintrust when evaluation belongs in CI and a platform fee is acceptable: the Eval API, the scorer library and the eval-action workflow are one integration. Choose Phoenix when traces cannot leave your infrastructure, because Braintrust's BYOC and self-hosted deployments require an Enterprise contract while Phoenix is free to self-host under the Elastic License 2.0.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.