Tools/LLMOps & evals

Langfuse review: tracing, prompts and evals you can host yourself

Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.

Type
LLM observability
Pricing
MIT · paid from $59 per month

··10 min read

  • LLM observability
  • Tracing
  • OpenTelemetry
  • Self-hosting
  • Evaluation
A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.

Key takeaways

  • Langfuse is the most complete open-source option for tracing, prompt management and experiments in one backend, and the MIT licence covers the whole core rather than a sample.
  • A production self-hosted deployment is two application containers plus PostgreSQL, ClickHouse, Redis and object storage, which is a bigger commitment than the Docker Compose page implies.
  • Cloud bills data points, not seats: one agent turn with three observations and two scores is six units, and the pricing page lists a free Hobby plan and Core at $29 a month for 100,000 units.
  • Instrumentation is queued and batched in the background, so tracing adds no request latency, but a process that exits without flushing loses the tail of its traces.
  • Langfuse Cloud stops serving v3 endpoints on 16 November 2026, so an existing v3 integration needs a migration window rather than an upgrade at leisure.

Langfuse is an open-source platform for tracing, prompt management and evaluation of LLM applications, and it is the most complete one of those that ships under a permissive licence. The verdict up front: teams that want tracing, prompts and experiments in one backend, and that can look after a ClickHouse deployment, should take it. Teams that want a weekend of setup should not.

It sits in the same layer as LangSmith, Arize Phoenix and Helicone, and it competes on two axes that matter: whether the platform can run inside your own network, and whether you instrument your code or your model provider. Langfuse answers yes on the first question for the core product, and answers both ways on the second: drop-in wrappers for the common SDKs, a context manager for everything else, and a plain OpenTelemetry endpoint for code that is neither Python nor JavaScript. The wider case for tracing agents is in agent observability with OpenTelemetry.

What it is

Langfuse is four products in one deployment: a trace store for LLM calls, a prompt registry with versioning and a playground, an experiment runner over datasets, and a dashboard layer for cost, latency and scores. All four run on the same code whether hosted or self-hosted, and the self-hosted edition is not a reduced build.

  • MIT licence for the core. The repository carries a ClickHouse Inc copyright, and only the directories under ee/ sit under a separate commercial licence.
  • Self-hosted on Docker Compose, Helm or Terraform templates for AWS, Azure and GCP, running the same containers as Langfuse Cloud.
  • Four storage services in a normal deployment: PostgreSQL, ClickHouse, Redis or Valkey, and S3-compatible object storage.
  • Python and JavaScript SDKs, plus a documented OpenTelemetry ingestion endpoint with a version header on every request.
  • 35,468 stars and 3,940 forks on GitHub in October 2026, with release v4.54.0 shipped on 7 October 2026.
  • Owned by ClickHouse since January 2026, which is also the database it stores traces in.
  • Cloud regions in the EU, US and Japan, plus a HIPAA region on the enterprise plan.

How it works

The ingestion path explains both the operational cost and the latency story. An SDK or a collector sends a batch of events; the web container writes that batch straight to object storage and leaves only a reference in Redis; a worker picks the events up and writes them into ClickHouse, where traces, observations and scores live. PostgreSQL holds the transactional side: projects, API keys, prompt versions.

How a trace reaches storage in LangfuseThe application or an OpenTelemetry collector sends a batch of events to the Langfuse web container. The web container writes the batch to object storage and leaves a reference in Redis. A worker reads the batch from object storage and writes traces, observations and scores into ClickHouse. PostgreSQL holds the transactional data such as projects and prompt versions, and Redis holds the queue and the API key and prompt caches.Trace ingestionsame stack in cloud and self-hostedApp or SDKbatched eventsLangfuse WebUI and APIS3 or blobraw eventsWorkerasync ingestRedisqueue, key cachePostgreSQLprojects, promptsClickHousetraces, scoresReads hit ClickHouse for traces and PostgreSQL for project and prompt data.Raw events reach object storage first: an outage delays data, it does not lose it
Every event is persisted to object storage before it reaches the analytical database.

Two consequences follow. A ClickHouse outage does not lose events, because the raw batch is already in object storage and is replayed. And trace queries never touch PostgreSQL, which is why the ClickHouse schema is shaped as a wide, mostly immutable observations table. The same design is what lets the vendor claim more than 90 billion observations a month across the platform; that is a company figure on its own infrastructure, not an independent benchmark.

Getting started

Two entry points cover most cases. The OpenAI wrapper records calls without changing the call sites, and the context manager is for code that does not use the OpenAI SDK or where one span should wrap several model calls. Both read credentials from the environment, so the same instrumentation runs against the EU, US, Japan or HIPAA cloud region, or against a self-hosted deployment.

import os

from langfuse import get_client
from langfuse.openai import openai

os.environ["LANGFUSE_PUBLIC_KEY"] = "pk-lf-..."
os.environ["LANGFUSE_SECRET_KEY"] = "sk-lf-..."
os.environ["LANGFUSE_HOST"] = "https://cloud.langfuse.com"  # EU region

langfuse = get_client()

# 1. Drop-in: every OpenAI call is recorded as a generation.
openai.chat.completions.create(
    model="gpt-4o",
    name="calculator",
    messages=[{"role": "user", "content": "1 + 1 = "}],
)

# 2. Explicit: a span around one unit of work, a generation inside it.
with langfuse.start_as_current_observation(as_type="span", name="answer-question") as span:
    span.update(input={"question": "What is the capital of France?"})
    with langfuse.start_as_current_observation(
        as_type="generation", name="llm-call", model="gpt-4o"
    ) as generation:
        generation.update(output="Paris.")
    span.update(output="answered")

langfuse.flush()  # required in short-lived processes

Anything that already emits OTLP spans can skip the SDKs entirely: the ingestion endpoint accepts a standard trace payload with the project keys as basic auth and the ingestion version in a header. That path is what makes Langfuse a defensible choice for a polyglot estate, where Python instrumentation would be one more service to maintain.

Frameworks get first-class integrations rather than wrappers: the docs list OpenTelemetry, the Vercel AI SDK, LangChain for Python and JavaScript, LlamaIndex, CrewAI, AutoGen, Google ADK and Ollama, plus proxy-based logging for teams whose calls already go through LiteLLM. That last entry settles the architecture question, because a team routing every model call through a gateway can take its traces from the gateway instead of from the application.

Prompts and evaluations

Tracing is the part every competitor has. What decides whether Langfuse earns its heavier deployment is that prompts and evaluations sit on the same traces: a prompt is versioned in the UI, fetched by the SDK with a client-side cache and released through labels, while datasets and experiments run that prompt against stored inputs and write the scores back onto the observations.

  • Prompt versioning with release management and composability, with protected deployment labels on the enterprise plan.
  • Client-side prompt caching in the SDKs, revalidated in the background, with a read-through cache in Redis on the server.
  • LLM-as-a-judge evaluators, custom scores from code, user feedback capture and annotation queues in the UI.
  • Code evaluators that run deterministic Python or TypeScript checks on live observations, added in v4.
  • Monitors that watch cost, latency and quality thresholds and notify through Slack, webhooks or GitHub Actions.

The opinionated part: this is the right place to run offline evaluation. Because an experiment writes its scores onto the same observation identifiers the production trace uses, a regression found in a dataset run stays traceable to the request shape that produced it, which is the step most eval tooling leaves to a spreadsheet. The price is a prompt release process you now own, and a prompt registry that nobody updates is worse than no registry at all.

What it costs to run

Self-hosting is free with no usage limit, which makes the cloud price the only thing left to model. Cloud bills data points, not seats: a unit is any trace, observation or score you send, so one agent turn with three observations and two scores is six units.

  • PostgreSQL for transactional data, ClickHouse for traces, observations and scores, Redis for the queue and the API key and prompt caches.
  • Events land in object storage before the database, so an analytics outage delays data instead of losing it.
  • Background migrations move long-running schema work off the upgrade path, which shortens upgrade downtime.
  • Client-side data masking ships in the open-source edition; server-side masking, audit logs, SCIM, project-level RBAC and retention policies need an enterprise licence key.
  • The self-hosted enterprise edition is bundled with ClickHouse Cloud, BYOC or Private, and its price is additive to that ClickHouse plan.
  • Templates exist for Kubernetes, AWS, Azure and GCP; Render and Railway are community-supported only.

The operational judgement: this is a platform, not a sidecar. Running it means watching ClickHouse, which means someone has to know ClickHouse. Teams that already run it gain a capability they could not otherwise buy; teams that do not should start on the free Hobby plan and keep self-hosting for the day the data-residency question becomes real.

PlanPriceIncludedWhat changes
HobbyFree50,000 units a month30 days of data, two users, 1,000 ingestion requests a minute
Core$29 a month100,000 units90 days of data, unlimited users, 4,000 requests a minute
Pro$199 a month100,000 unitsthree years of data, 20,000 requests a minute, SOC 2 and ISO 27001 reports
Enterprise$2,499 a month100,000 unitsaudit logs, SCIM, custom rate limits, uptime SLA

Usage above the included units is billed at graduated rates: $8 per 100,000 units up to one million, $7 up to ten million, $6.50 up to fifty million and $6 beyond. The pricing page puts a one-million-unit month on Core at $101 and a twenty-five-million-unit month on Pro with the Teams add-on at $2,176. Model your unit count before comparing this with a per-seat or per-gigabyte price, because the same workload can differ by an order of magnitude depending on how many observations you emit per request.

Where it falls short

The weaknesses first, because they are the reasons to buy something else. Instrumentation still has to be written, and an application that calls three providers through a gateway needs three integrations or one OpenTelemetry pipeline. The interface is dense, and the distance from a trace to a dashboard that answers a question is measured in days. Unit pricing punishes verbose instrumentation, so the cheapest way to cut a bill is to emit less detail, which is exactly the wrong instinct during an incident. And the product now belongs to a database vendor: ClickHouse bought Langfuse in January 2026, which has clearly helped the roadmap and also means the commercial enterprise edition is increasingly a ClickHouse conversation.

ToolLicenceWhere it runsStrongest at
LangfuseMIT core, ee/ directories under a commercial keyCloud, or your own Docker, Helm or Terraform deploymenttraces, prompts and experiments on one backend
LangSmithclient SDK on MIT, the platform is hostedLangChain's hosted serviceteams that have already chosen LangGraph
Arize PhoenixElastic License 2.0Self-hosted, tracing through OpenInference and OpenTelemetryevaluation-first workflows over your own traces
HeliconeApache-2.0Self-hosted or cloudcapturing every call with no SDK change, through a proxy

The comparison that matters is not feature count. Phoenix is free to run and stricter about the evaluation story, but Elastic License 2.0 is source-available rather than open source and stops you offering it as a service. Helicone is the cheapest route to cost and latency numbers on every call, at the price of a proxy in the request path. LangSmith is the smoothest option once the framework decision is made, and the one whose price scales with the size of the engineering team.

Verdict

Langfuse is the tool this category needed: a permissively licensed platform that does not ask you to change frameworks, put a proxy in the request path, or give up the raw data. Take a position on it: it is the right default for a team running several agent surfaces against a real bill, and the wrong choice for a single prompt in a side project.

  1. Adopt it if traces, prompt versioning and experiments have to live in one place and someone can operate ClickHouse.
  2. Adopt it if prompt versions must be auditable, which is the usual requirement in a procurement conversation in the EU.
  3. Adopt it if you already emit OpenTelemetry and would rather have one backend than per-framework instrumentation.
  4. Do not adopt it if the application makes a handful of model calls a day; a free tier plus structured logs costs less attention.
  5. Do not adopt it if nobody will tune instrumentation. The bill scales with the number of observations, not with the number of users, and the cheapest saving is to turn off the detail you need at 3am.

Sources

  • taghrefchildren
  • taghrefchildren
  • taghrefchildren
  • taghrefchildren
  • taghrefchildren
  • taghrefchildren
  • taghrefchildren
  • taghrefchildren

Frequently asked questions

How much does Langfuse cost?

Self-hosting is free with no usage limit under the MIT licence. Langfuse Cloud has a free Hobby plan with 50,000 units a month, 30 days of data and two users, Core at $29 a month with 100,000 units and 90 days of data, Pro at $199 and Enterprise at $2,499. Usage above the included units runs from $8 per 100,000 units down to $6 at very high volume.

What counts as a billable unit in Langfuse?

Any tracing data point you send to the platform: a trace, an observation inside it such as a span, event or generation, or a score. A single agent turn with three observations and two evaluation scores therefore costs six units, which is why instrumentation granularity matters more than the number of users.

Does Langfuse add latency to my application?

The docs say no: the SDKs queue trace events locally and flush them in batches in the background, so the request path is not blocked. In short-lived processes such as Lambda handlers, cron jobs and CLIs you must call flush() at the end, or the last observations in the batch are lost when the process exits.

Langfuse or LangSmith?

Choose Langfuse if you want an MIT-licensed platform you can run in your own VPC and traces, prompts and experiments in one model. Choose LangSmith if you build on LangGraph and want first-party tracing with annotation queues and nothing to assemble. The cost models differ too: Langfuse charges per data point with unlimited users, LangSmith charges per seat plus traces.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.