Tools/LLMOps & evals

Helicone: observability that sits in the request path

Helicone is an Apache-2.0 LLM gateway and observability platform. What the proxy architecture buys, what it costs you, and how the self-hosted stack really looks.

Type
LLM observability
Pricing
Open core · from $20 per month

··10 min read

  • LLM observability
  • OpenTelemetry
  • Cost tracking
  • Proxy
  • Gateway
Cover: Helicone, an LLM observability platform, showing the request path from application through the edge proxy to the provider and back into the log store

Key takeaways

  • Helicone is Apache-2.0 with about 6,200 GitHub stars and covers 100 plus models behind one OpenAI-compatible endpoint.
  • Integration is one line: a different base URL, or the OpenLLMetry async path if the proxy must stay off the critical path.
  • The proxy adds caching, rate limits, fallbacks and retries, and the async path gives all of those up.
  • The vendor benchmark reports a mean of 2.21 seconds both direct and proxied on 500 interleaved requests, measured on text-ada-001.
  • Self-hosting runs five components — web, worker, Jawn, Supabase and ClickHouse plus MinIO — and Jawn no longer proxies, so the gateway is a separate deployment.

Helicone is an Apache-2.0 licensed LLM gateway and observability platform, and the review position is that the architecture is the whole story: Helicone works by sitting between the application and the model provider, which buys an enormous amount of functionality for one changed line of code and costs a network hop, a data-residency question and a vendor on the critical path. For teams that want a dashboard on Monday, that trade is good. For teams with a latency budget measured in milliseconds or a rule that prompts cannot leave the network, it is the wrong shape.

It competes in two directions at once. Against pure tracing backends such as Langfuse, Arize Phoenix and LangSmith, it wins on time-to-first-dashboard and on the gateway features bolted to the proxy, and loses on OpenTelemetry-native instrumentation and on evaluation tooling. Against API routers such as LiteLLM or OpenRouter, it is observability first with routing attached.

What it is

Two products share one codebase. The gateway is an OpenAI-compatible endpoint in front of 100 plus providers, reached by pointing the base URL at ai-gateway.helicone.ai. The observability platform is what records the requests that pass through, stores them in ClickHouse, and answers questions about cost, latency, sessions and users through a dashboard, a SQL-like query language called HQL, alerts, reports and webhooks.

  • Licence Apache 2.0, about 6,200 GitHub stars, and a note that Helicone joined Mintlify in 2025.
  • One-line integration: change the base URL, or send an async log through the OpenLLMetry instrumentation.
  • Cost tracking computed from token usage and a public price database covering more than 300 models and providers.
  • Gateway features on the request path: edge cache, custom rate limits by request count, cost or property, and automatic fallbacks.
  • Operational surface: sessions, per-user metrics, custom properties, scores, datasets, webhooks and an MCP server over the data.
  • Self-hosting that runs five services: a web frontend, a Cloudflare Worker, the Jawn collector, Supabase and ClickHouse, plus MinIO for bodies.

How it works

The request path is deliberately thin. Unless a header enables a feature, the worker forwards the request and returns the response untouched; after the response is complete, the proxy ships logs to Kafka for a separate service to consume. The availability documentation states the design intent directly: all business logic falls back to plain proxying on any error, so a bug in observability degrades to a working relay.

The Helicone request pathTop row: the application sends a chat completion through the OpenAI SDK to the Helicone gateway, which sits on Cloudflare Workers and either serves a cache hit or forwards to the provider and returns the response. Bottom row: after the response is complete, logs go to Kafka and then into ClickHouse for analytics and MinIO for request and response bodies, where the dashboard, HQL, alerts and reports read from. Below them a bar: the async path bypasses the gateway entirely, logging after the response from the application itself, and gives up caching, rate limits and retries in exchange.The Helicone request pathproxy first, logs secondYour appOpenAI SDK, one linerequestGatewayCloudflare Workers, edgeforwardProvider100+ models, one APIresponseafter the response onlylogs to Kafkaedge cacherate limits, fallbacksClickHouse for analyticsMinIO for request bodiesDashboard, HQL, alerts, reports, MCPAsync path skips the gateway: logs after the response, loses caching, rate limits and retriesOn error the worker falls back to plain proxying, so observability degrades to a relay
The gateway is on the critical path only in the proxy mode, and the logging happens after the response in both.

One consequence is worth stating plainly: in proxy mode every prompt and every completion transits a third-party edge network by default. The self-hosted deployment removes that, but it also removes the part of the product that makes Helicone easy, because a self-hosted Jawn no longer proxies at all — the docs record that the gateway routes were removed and the AI Gateway must be deployed separately.

Getting started

The whole integration is a base URL. The example below also switches on the edge cache, which is the fastest way to see the platform do something visible: the second identical request is served from Cloudflare's KV store rather than from the provider.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://ai-gateway.helicone.ai",
    api_key=os.environ["HELICONE_API_KEY"],
    default_headers={
        "Helicone-Cache-Enabled": "true",
        "Cache-Control": "max-age=3600",
        "Helicone-Cache-Seed": "user-123",
        "Helicone-Property-Env": "production",
    },
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarise the refund policy."}],
)

print(response.choices[0].message.content)
# The next identical call returns Helicone-Cache: HIT and skips the provider.

The cache key is a hash of the cache seed, the request URL, the whole request body and the relevant headers, which makes it exact-match rather than semantic: change one word or one temperature and it misses. Durations run from an hour to a 365-day maximum, buckets up to 20 stored responses are available for non-deterministic prompts, and Helicone-Cache-Ignore-Keys excludes JSON fields such as a request id or a provider prompt_cache_key from the key.

Proxy or async, pick one deliberately

This is the decision that shapes everything else. Helicone documents both modes and is unusually honest about the gap between them, which makes the choice easy to reason about even though the answer is rarely comfortable.

CapabilityProxy modeAsync mode
On the critical pathyesno
Gateway cache and bucketsyesnot available
Custom rate limitsyesnot available
Automatic retriesyesnot available
Prompt auto-formattingyesnot available
Setup effortone base URLinstrument with OpenLLMetry
Sessions and user metricsyesyes
Custom properties and scoresyesyes
Data pathprompts cross Helicone's edgeprompts stay with the provider
Use it whenyou want caching and rate limitspropagation delay is unacceptable

The reasonable conclusion is that these are two products sharing a name. The proxy is a gateway with analytics attached; the async integration is a logging library with a web UI. Teams that take the proxy get caching, rate limiting and fallbacks that they would otherwise build, and pay for it with an extra hop and a data-residency decision. Teams that take the async path keep their latency budget and their network boundary, and build the gateway themselves or go without it.

Latency and the cost of a hop

The published benchmark is small but unusually well specified: 500 requests with unique prompts interleaved between OpenAI and Helicone inside the same one-second window, alternating which endpoint was called first, with the prompt context maximised, on text-ada-001, logging round-trip latency for both sets.

StatisticOpenAI direct (s)Helicone proxied (s)
Mean2.212.21
Median2.872.90
Standard deviation1.121.12
p903.273.29
Maximum3.563.76
Method500 interleaved requestssame prompts, alternating order
Modeltext-ada-001text-ada-001, max context
Overheadbaselinevisible only in the maximum
Who ran itvendorvendor

Read the maximum row carefully. A 200 millisecond difference on a three-and-a-half second response is 6 per cent, and the same absolute hop applied to a 300 millisecond model call or to a 40 millisecond tool call would dominate it. The benchmark is also a vendor test on a retired model, so it demonstrates that the overhead is bounded, not what it will be for a given workload.

Pricing and what the meter is

The free Hobby plan covers 10,000 requests a month, one seat, one organisation, one gigabyte and seven days of retention. Pro is $79 a month for unlimited seats, one organisation, alerts, reports, HQL, one month of retention and 1,000 ingested logs per minute. Team is $799 for five organisations, SOC 2 and HIPAA options, three months of retention and 15,000 logs per minute. Enterprise adds unlimited organisations, SAML SSO and on-prem deployment.

The number to watch is the one nobody prices prominently: ingestion. A Hobby workspace ingests 10 logs a minute, Pro 1,000, Team 15,000. A production application can emit several logs per user request — a model call, a retrieval step, a tool call — so a moderately busy product exhausts the Pro ceiling before it exhausts the request allowance, and the plan that fixes it costs 799 dollars a month. Storage is metered separately above the first gigabyte.

PlanPriceRequestsIngestionRetentionHobby$010,000 per month10 logs per minute7 daysPro$79 per monthusage-based1,000 logs per minute1 monthTeam$799 per monthusage-based15,000 logs per minute3 monthsEnterprisecustomusage-based30,000 logs per minuteforever
Seats1unlimitedunlimitedunlimited
Organisations115unlimited
Notable additionsnothingHQL, alerts, reportsSOC 2, HIPAASAML SSO, on-prem
Self-hostfreefreefreeHelm chart on request

Self-hosting the whole stack

The Apache-2.0 repository contains the entire platform, and the README is direct about the operational shape: five services, and a manual deployment that it marks as not recommended in favour of the Docker compose file or an Enterprise Helm chart.

  • Web: the dashboard frontend, a Next.js application.
  • Worker: the proxy, deployed as Cloudflare Workers in the hosted version.
  • Jawn: the log collector, an Express service with Tsoa-generated routes.
  • Supabase: the application database and authentication.
  • ClickHouse: the analytics store that answers cost and latency questions.
  • MinIO: object storage for request and response bodies.

Where it falls short

The honest weaknesses are structural, not cosmetic. Helicone's data model is built around requests that pass through its gateway, so a team that already traces with OpenTelemetry spans has to choose between two instrumentation stacks. The evaluation story is thin next to Langfuse or Braintrust: prompts, playground and scores exist, but there is no first-class experiments workflow. And the ownership question is live — Helicone joined Mintlify in 2025, which removes the open-core lock-in risk but adds the usual vendor dependency.

HeliconeLangfuseArize Phoenix
LicenceApache 2.0MIT core, enterprise extrasElastic License 2.0
Instrumentationproxy or OpenLLMetryOpenTelemetry nativeOpenInference on OTel
Strongest suitone-line start, gateway featuresevaluation and prompt workflowslocal notebook-style evaluation
Weakest suitnot OTel-native for tracingheavier to start than a base URLELv2 is not permissive
Gateway featurescache, limits, fallbacksnono
Self-hostApache 2.0, five servicesMIT, Docker Compose or HelmELv2, self-hostable
Cost trackingbuilt inconfigurable per modelseparate product
Best fitteams wanting visibility this weekteams running evaluationsteams tracing and evaluating locally

Langfuse is MIT-licensed, was acquired by ClickHouse in January 2026 with the licence and self-hosting explicitly unchanged, and stores traces in ClickHouse as well; if the deciding factor is a permissive licence with a real evaluation stack, it is the stronger tool. Arize Phoenix is built on OpenInference conventions and runs anywhere, including air-gapped, but its Elastic License 2.0 is not a permissive one. Helicone's argument is not capability, it is time: the distance between a repository and a cost dashboard measured in requests per user is one line of code.

Verdict

Helicone is the right tool when nobody can answer what the LLM spend was last Tuesday, and the wrong tool when latency, data residency or tracing standards decide the architecture. Its engineering is unremarkable in the best sense: the proxy is thin, logs go out after the response, and the product is Apache-2.0 from the gateway to the dashboard. Its weakness is equally plain: it wants to be your gateway, and everything that makes it comfortable is a consequence of that.

  1. Use it when the team needs per-request cost, latency and error visibility within a week and has no tracing stack yet.
  2. Use it when the edge cache, custom rate limits and automatic fallbacks are worth having on the request path.
  3. Use the async OpenLLMetry path when you want Helicone's analytics without a proxy hop, and accept losing the gateway features.
  4. Avoid it when prompts must not leave your network or when the proxy hop eats a hard latency budget.
  5. Avoid it when the platform already traces with OpenTelemetry and a second instrumentation model would fragment the data.
  6. Reconsider it above roughly 1,000 ingested logs a minute, or when the request allowance is large enough that the usage meter needs its own line in the budget.

Sources

  1. Helicone quickstart
  2. Helicone AI Gateway overview
  3. Helicone: latency impact and benchmark
  4. Helicone: proxy versus async integration
  5. Helicone: LLM caching
  6. Helicone: self-hosting with Docker
  7. Helicone: how we calculate cost
  8. Helicone pricing
  9. Helicone repository on GitHub

Frequently asked questions

Is Helicone open source?

Yes, under Apache 2.0. The repository at roughly 6,200 stars contains the gateway, the log collector and the dashboard. The hosted tiers add features such as SOC 2 reports, HIPAA options, SAML SSO and on-prem deployment terms.

Does the Helicone proxy add latency?

Helicone's own benchmark sends 500 interleaved requests to OpenAI directly and through Helicone and reports a mean of 2.21 seconds in both cases, with p90 at 3.27 versus 3.29 seconds. It is a vendor test on text-ada-001, so treat it as evidence that the overhead is small rather than as a number to plan against.

Can Helicone stay out of the critical path?

Yes, through the OpenLLMetry async integration, which logs after the response and never sits between the application and the provider. The documentation is explicit about the trade: the async path loses bucket caching, custom rate limits and retries.

How does Helicone work with several providers?

The AI Gateway exposes 100 plus models through one OpenAI-compatible endpoint and translates the request to each provider's format. With credits, Helicone holds the provider keys and claims zero markup; you can also bring your own keys.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.