Tools/LLMOps & evals

LangSmith: an observability product that grew into an agent platform

LangSmith review: per-trace billing, 14 and 180 day retention, OpenTelemetry ingestion, offline and online evals, and where the platform's pull towards an agent runtime shows up.

Type
LLM observability
Pricing
Developer free · from $39 per month

··10 min read

  • Tracing
  • LLM evals
  • OpenTelemetry
  • Datasets
  • Retries
A LangSmith trace path: traced application code posts runs through an SDK background thread into an ingest queue, which writes traces to ClickHouse for the dashboard and the API.

Key takeaways

  • The billable unit is the trace, not the span or the token. Developer includes 5,000 base traces a month for one seat, Plus costs $39 a seat with 10,000 included, and the Developer plan without a payment method on file is capped at 5,000 traces a month.
  • Base traces are retained for 14 days and extended traces for 180, and online evaluators, automation rules and API feedback with extend_trace_retention upgrade traces to the more expensive tier unless you opt out.
  • Ingest limits are hourly and per plan: 50,000 events and 500 MB an hour on Developer without payment details, 250,000 events and 2.5 GB with them, and 500,000 events and 5 GB on Plus.
  • The rest of the platform is billed in LangChain Standard Units at $1.00 each, including Engine, Fleet, Sandboxes and the LLM Gateway, which makes a seat price a poor guide to the invoice.
  • OpenTelemetry ingestion works through the SDK or an OTLP endpoint, but a span whose parent never arrives is buffered and then silently dropped, which is the failure mode to watch in a partial fan-out.

LangSmith is a hosted tracing and evaluation platform for LLM applications: every call your agent makes becomes a tree of runs with inputs, outputs, latency and token counts that you can search, score and compare. The position after reading the documentation: it is the most capable tool of its kind, and the most expensive one to understand, because it is no longer only an observability product.

It competes with Langfuse, Arize Phoenix and Helicone for the observability budget, and increasingly with its own runtime, because the same vendor now sells Deployment, Studio, the Engine, Fleet, Sandboxes and an LLM Gateway on the same pricing page. That expansion is the single most important thing to weigh, since each extra service is metered separately.

What it is

The platform has four parts that share one bill: tracing, evaluation, prompt management and an agent runtime. The resource model matters more than the features. An organisation holds workspaces, workspaces hold tracing projects, datasets, annotation queues and prompts, and every trace lives in one project. The core facts:

  • A trace is one execution, made of nested runs; a run is created and then updated as the work progresses, which is why an event limit and a trace limit are different numbers.
  • Tracing works through the langsmith SDK decorators, a REST ingest API or OpenTelemetry spans from any instrumented application.
  • Offline evals run against datasets of examples with reference outputs; online evals score live traces without references.
  • Evaluators can be code, LLM-as-judge, a typed decision model, pairwise or human, and one evaluator can be attached to several projects.
  • Annotation queues, dataset versions and splits, a prompt registry with commit tags and a playground are all included.
  • Self-hosted deployment exists, but only as an Enterprise add-on behind a licence key.

How it works

Instrumentation runs in your process and posts runs to LangSmith over HTTPS. The SDK sends from a background thread and batches up to 100 runs from one session into a single API call, so tracing does not sit in the request path. A server-side queue then handles ingestion, retries and integrity checks before writing into the trace store.

The path of a LangSmith traceTraced application code creates runs. The SDK posts them in batches from a background thread, because a rate limit stops the first 5000 posts to the runs endpoint in a minute. An ingest queue retries and stores the runs in ClickHouse. Dashboards, monitors and the query API read from that store.one request, one trace treeyour apptraceableSDK threadbatched postsqueueretrytrace storeClickHouse429 after 5,000 posts per minutedashboards, monitors and the query API all read from the same store
Because the queue is asynchronous, a trace that is accepted with a 200 can still fail to arrive, and a short-lived process can exit before its runs are posted.

That queue is why the ingestion limits are expressed as a fixed window rather than a smooth average, and why a 429 is a normal event to handle rather than an outage. The load balancer enforces fixed per-minute caps on every plan: 5000 POST or PATCH requests to the runs endpoints, 5000 to feedbacks, 2000 for any other endpoint, and 30 deletes. The SDK batches, which is what keeps a busy application under those numbers.

Getting started

Two environment variables switch tracing on without touching code, which matters for local development: LANGSMITH_TRACING gates the decorator and the context manager, and LANGSMITH_PROJECT names the destination project, defaulting to default. The minimal Python setup traces a pipeline as nested runs:

import asyncio

from langsmith import Client, traceable
from openai import AsyncOpenAI

client = Client()
llm = AsyncOpenAI()


@traceable(run_type="retriever", name="retrieve_docs")
async def retrieve_docs(question: str) -> list[str]:
    return ["Annual report: revenue up 12 percent."]


@traceable(run_type="llm", name="answer")
async def answer(question: str, context: list[str]) -> str:
    reply = await llm.chat.completions.create(
        model="gpt-5.4-mini",
        messages=[{"role": "user", "content": f"{question}\n{chr(10).join(context)}"}],
    )
    return reply.choices[0].message.content


@traceable(name="support_agent")
async def support_agent(question: str) -> str:
    return await answer(question, await retrieve_docs(question))


async def main() -> None:
    try:
        print(await support_agent("How did revenue move?"))
    finally:
        await client.flush()  # background thread must finish before exit


asyncio.run(main())

The decorator propagates context, so the three functions appear as a tree without any manual parent wiring, and run_type decides how the dashboard renders a node: llm gives token counts and latency, retriever and tool mark the other kinds of step. The same client runs offline evaluations against a dataset, which is where the value starts to compound.

The same API runs the eval loop. evaluate takes a target function, a dataset and a list of evaluators, produces an experiment with a run per example, and can be driven from CI: Evaluate each committed prompt change against the tagged dataset version and fail the build if the groundedness score drops by more than a point. LangSmith versions datasets automatically when examples change, so a tag can pin a CI run to one state of the data. The documentation is blunt about the starting point: write five to ten curated examples of good output before writing any evaluator.

What it costs

Seats are the visible price and traces are the metered one. The Developer plan is free for a single seat with 5,000 base traces a month included; Plus is $39 per seat a month with 10,000 included and unlimited extra seats; Enterprise is quoted and adds self-hosted and hybrid deployment, custom SSO, and attribute-based and role-based access control.

PlanSeatsTraces includedIngest ceiling per hourRetention
Developer, no payment details15,000 a month, and a monthly cap of 5,00050,000 events and 500 MB14 days, extended by upgrade
Developer with payment details15,000 a month250,000 events and 2.5 GB14 days, extended by upgrade
PlusUnlimited, $39 each10,000 a month500,000 events and 5 GB14 or 180 days
EnterpriseCustomCustomCustomUp to 180 days, configurable

An event is the creation or the update of a run, so a run created and then patched inside the same clock hour counts twice against the hourly limit, and a 2 MB run later updated to 3 MB counts 5 MB against the ingest volume limit. That is the mechanic to model in a capacity plan: trace shape, not request count, drives the ceiling.

The rest of the platform is billed in LangChain Standard Units, with one LSU priced at $1.00. The Engine is scheduled every six hours and a single run is estimated at 7 to 45 LSU depending on trace volume and issue count. Fleet includes 7 LSU on Developer and 37 LSU on Plus, Sandboxes 8 LSU, and a perceived-error evaluator run 0.015 LSU. A $39 seat therefore says very little about the invoice once the runtime is in use.

Native tracing or OpenTelemetry

LangSmith ingests OpenTelemetry spans two ways. With the SDK integration, LANGSMITH_OTEL_ENABLED=true makes LangChain and LangGraph emit spans through the LangSmith exporter, and LANGSMITH_OTEL_ONLY=true stops it sending to LangSmith's own format as well. With any other application, point a standard OTLP exporter at the base endpoint https://api.smith.langchain.com/otel; regional endpoints exist for EU, APAC and AWS US. The exporter appends the signal path itself, so putting /v1/traces in the base URL gives a 404.

pip install "langsmith[otel]"          # needs langsmith >= 0.3.18, 0.4.25 recommended

export LANGSMITH_TRACING=true
export LANGSMITH_OTEL_ENABLED=true
export LANGSMITH_ENDPOINT=https://api.smith.langchain.com
export LANGSMITH_API_KEY=...

# fan out one OTLP stream to LangSmith and to the rest of the stack
export OTEL_EXPORTER_OTLP_ENDPOINT=https://api.smith.langchain.com/otel
export OTEL_EXPORTER_OTLP_HEADERS="x-api-key=...,Langsmith-Project=support"

# OTel-only, for teams that do not want a second transport
export LANGSMITH_OTEL_ONLY=true

The trade is between convenience and overhead. The OTel path costs an attribute-mapping exercise, because span attributes have to be labelled with the langsmith namespace to become run types, run IDs and dotted order; LangChain's own announcement describes the OpenTelemetry route as having slightly higher overhead and recommends the native format when LangSmith is the only destination. The native format also gives pending runs that appear in the UI while the work is still going.

Where it shingle

The weaknesses first, because they decide whether you need this product. The billing unit is the trace, which quietly rewards sampling and punishes an application that traces every request; the load-balancer caps are per service key or personal access token, not per organisation, so a horizontally scaled fleet needs either more keys or the SDK's batching. Workspace role-based access control is Enterprise only, so a growing team on Plus shares one role model. And the resource hierarchy is being rewritten underneath you: workspaces were tenants, agents are in beta as a grouping above projects, and an agent-based workspace addresses traces by agent and environment rather than by project.

ToolPrimary strengthHostingWhat it gives up
LangSmithTracing plus evals plus a managed agent runtimeCloud, or self-hosted on EnterpriseOpen core is not an option: the SDK is free, the platform is not
LangfuseMIT core you can self-host and inspectCloud or self-hostLess of the managed runtime around the traces
Arize PhoenixApache-2.0, built on OpenTelemetrySelf-host or cloudA narrower product surface around datasets and evals
HeliconeCheap request-level logging with fast setupCloud or proxy deploymentLess depth on run trees and evaluation workflows

What LangSmith does not give up is measurement: it is the reference implementation for run trees, thread-level conversation views and dataset-driven regression testing, and the annotation queue with reservations is a genuine answer to the question of who labels what. If your application already runs on LangGraph, choosing it is the obvious path and the switching cost is the runtime, not the traces.

Verdict

LangSmith is worth paying for when an agent's behaviour, not its uptime, is the thing that breaks, and when you need humans in the loop labelling runs and datasets that drive a release gate. It is worth less than its price when you only want a request log, when one trace per user request is too many traces for the budget, or when the platform sprawl makes it unclear what you are paying for.

  1. Adopt it if you need dataset-driven regression tests that run in CI, not just a trace viewer.
  2. Adopt it if several people have to label runs, because annotation queues with reservations and dataset export are hard to rebuild.
  3. Adopt it if you already build on LangGraph, and factor the runtime into the decision rather than pretending the traces are separate.
  4. Be careful if you plan to trace every request: model the event and byte ceilings per hour before you turn instrumentation on in production.
  5. Look elsewhere if you need self-hosting without a sales conversation, or workspace-level access control, or if a flat price per trace fits your workload better than an LSU-metered platform.

Sources

  1. LangSmith pricing
  2. LangSmith documentation: administration overview
  3. LangSmith documentation: custom instrumentation
  4. LangSmith documentation: trace with OpenTelemetry
  5. LangSmith documentation: evaluation concepts
  6. LangSmith documentation: self-hosted deployment
  7. LangChain: end-to-end OpenTelemetry support in LangSmith

Frequently asked questions

How much does LangSmith cost for a small team?

The Developer plan is free for one seat with 5,000 base traces a month included, and without a payment method on file it is also capped at 5,000 traces a month. Plus is $39 per seat a month with 10,000 base traces included and unlimited extra seats. Enterprise is quoted, and it is the only tier that offers workspace role-based access control and self-hosted deployment.

Do I need LangChain to use LangSmith?

No. The langsmith SDK has a traceable decorator for Python, TypeScript, Kotlin and Java, plus a low-level RunTree API and a REST ingest path, so any code can be traced. LangChain and LangGraph just get the instrumentation for free. LangSmith also ingests OpenTelemetry spans, which the docs recommend for applications that already emit OTLP.

What happens to my traces after two weeks?

Base retention is 14 days, after which traces are no longer reachable in the UI or the API and the associated inputs and outputs are deleted within a day, while some trace metadata is kept for analytics and billing. Extended retention is 180 days and costs more; since 14 September 2026 that 180 days is the maximum for SaaS customers.

Can LangSmith run inside my own infrastructure?

Yes, but only as an Enterprise add-on with a licence key. A self-hosted instance runs the frontend, backend, platform backend, playground, queue and a code-execution service on top of ClickHouse for traces, PostgreSQL for operational data and Redis or Valkey for queues, with optional blob storage. LangChain recommends external database services in production rather than the bundled ones.

Will LangSmith train on my data?

No. The pricing FAQ states that LangSmith does not use your data to train models and that traces, prompts and outputs stay private to your organisation. That is a contractual answer rather than a technical control, so a self-hosted deployment is the option for teams who need the data inside their own perimeter.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.