[{"data":1,"prerenderedAt":813},["ShallowReactive",2],{"tool-helicone-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":28,"cover":56,"og":57,"expertise":58,"locales":59,"lang":60,"title":63,"description":64,"coverAlt":65,"url":66,"pricing":67,"kind":9,"metaTitle":68,"takeaways":69,"faq":75,"toc":88,"blocks":119,"others":580},"helicone","2026-07-03",10,"llmops",[9,10,11,12,13],"LLM observability","OpenTelemetry","Cost tracking","Proxy","Gateway",[15,16,17,18,19,20,21],"helicone llm observability","helicone vs langfuse","llm cost tracking tool","open source llm observability","llm gateway proxy","self-host llm observability","helicone cache",[23,25],{"name":10,"url":24},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":26,"url":27},"Apache License 2.0","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_License",[29,32,35,38,41,44,47,50,53],{"title":30,"url":31},"Helicone quickstart","https:\u002F\u002Fdocs.helicone.ai\u002Fgetting-started\u002Fquick-start",{"title":33,"url":34},"Helicone AI Gateway overview","https:\u002F\u002Fdocs.helicone.ai\u002Fgateway\u002Foverview",{"title":36,"url":37},"Helicone: latency impact and benchmark","https:\u002F\u002Fdocs.helicone.ai\u002Freferences\u002Flatency-affect",{"title":39,"url":40},"Helicone: proxy versus async integration","https:\u002F\u002Fdocs.helicone.ai\u002Freferences\u002Fproxy-vs-async",{"title":42,"url":43},"Helicone: LLM caching","https:\u002F\u002Fdocs.helicone.ai\u002Ffeatures\u002Fadvanced-usage\u002Fcaching",{"title":45,"url":46},"Helicone: self-hosting with Docker","https:\u002F\u002Fdocs.helicone.ai\u002Fgetting-started\u002Fself-deploy-docker",{"title":48,"url":49},"Helicone: how we calculate cost","https:\u002F\u002Fdocs.helicone.ai\u002Ffaq\u002Fhow-we-calculate-cost",{"title":51,"url":52},"Helicone pricing","https:\u002F\u002Fwww.helicone.ai\u002Fpricing",{"title":54,"url":55},"Helicone repository on GitHub","https:\u002F\u002Fgithub.com\u002FHelicone\u002Fhelicone","\u002Fimages\u002Fblog\u002Fhelicone\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fhelicone\u002Fog.jpg","ai-engineer",[60,61,62],"en","de","hu","Helicone: observability that sits in the request path","Helicone is an Apache-2.0 LLM gateway and observability platform. What the proxy architecture buys, what it costs you, and how the self-hosted stack really looks.","Cover: Helicone, an LLM observability platform, showing the request path from application through the edge proxy to the provider and back into the log store","https:\u002F\u002Fwww.helicone.ai","Open core · from $20 per month","Helicone: LLM observability with a proxy in the path",[70,71,72,73,74],"Helicone is Apache-2.0 with about 6,200 GitHub stars and covers 100 plus models behind one OpenAI-compatible endpoint.","Integration is one line: a different base URL, or the OpenLLMetry async path if the proxy must stay off the critical path.","The proxy adds caching, rate limits, fallbacks and retries, and the async path gives all of those up.","The vendor benchmark reports a mean of 2.21 seconds both direct and proxied on 500 interleaved requests, measured on text-ada-001.","Self-hosting runs five components — web, worker, Jawn, Supabase and ClickHouse plus MinIO — and Jawn no longer proxies, so the gateway is a separate deployment.",[76,79,82,85],{"q":77,"a":78},"Is Helicone open source?","Yes, under Apache 2.0. The repository at roughly 6,200 stars contains the gateway, the log collector and the dashboard. The hosted tiers add features such as SOC 2 reports, HIPAA options, SAML SSO and on-prem deployment terms.",{"q":80,"a":81},"Does the Helicone proxy add latency?","Helicone's own benchmark sends 500 interleaved requests to OpenAI directly and through Helicone and reports a mean of 2.21 seconds in both cases, with p90 at 3.27 versus 3.29 seconds. It is a vendor test on text-ada-001, so treat it as evidence that the overhead is small rather than as a number to plan against.",{"q":83,"a":84},"Can Helicone stay out of the critical path?","Yes, through the OpenLLMetry async integration, which logs after the response and never sits between the application and the provider. The documentation is explicit about the trade: the async path loses bucket caching, custom rate limits and retries.",{"q":86,"a":87},"How does Helicone work with several providers?","The AI Gateway exposes 100 plus models through one OpenAI-compatible endpoint and translates the request to each provider's format. With credits, Helicone holds the provider keys and claims zero markup; you can also bring your own keys.",[89,92,95,98,101,104,107,110,113,116],{"id":90,"title":91},"what-it-is","What it is",{"id":93,"title":94},"how-it-works","How it works",{"id":96,"title":97},"getting-started","Getting started",{"id":99,"title":100},"proxy-or-async","Proxy or async, pick one deliberately",{"id":102,"title":103},"performance","Latency and the cost of a hop",{"id":105,"title":106},"pricing","Pricing and what the meter is",{"id":108,"title":109},"self-hosting","Self-hosting the whole stack",{"id":111,"title":112},"where-it-shingles","Where it falls short",{"id":114,"title":115},"verdict","Verdict",{"id":117,"title":118},"sources","Sources",[120,124,127,130,133,149,150,153,162,165,166,169,172,175,182,183,186,255,258,259,262,331,334,340,341,344,347,415,416,419,433,438,439,442,521,524,525,528,543,549,550],{"type":121,"content":122},"paragraph",[123],"Helicone is an Apache-2.0 licensed LLM gateway and observability platform, and the review position is that the architecture is the whole story: Helicone works by sitting between the application and the model provider, which buys an enormous amount of functionality for one changed line of code and costs a network hop, a data-residency question and a vendor on the critical path. For teams that want a dashboard on Monday, that trade is good. For teams with a latency budget measured in milliseconds or a rule that prompts cannot leave the network, it is the wrong shape.",{"type":121,"content":125},[126],"It competes in two directions at once. Against pure tracing backends such as Langfuse, Arize Phoenix and LangSmith, it wins on time-to-first-dashboard and on the gateway features bolted to the proxy, and loses on OpenTelemetry-native instrumentation and on evaluation tooling. Against API routers such as LiteLLM or OpenRouter, it is observability first with routing attached.",{"type":128,"level":129,"id":90,"text":91},"heading",2,{"type":121,"content":131},[132],"Two products share one codebase. The gateway is an OpenAI-compatible endpoint in front of 100 plus providers, reached by pointing the base URL at ai-gateway.helicone.ai. The observability platform is what records the requests that pass through, stores them in ClickHouse, and answers questions about cost, latency, sessions and users through a dashboard, a SQL-like query language called HQL, alerts, reports and webhooks.",{"type":134,"ordered":135,"items":136},"list",false,[137,139,141,143,145,147],[138],"Licence Apache 2.0, about 6,200 GitHub stars, and a note that Helicone joined Mintlify in 2025.",[140],"One-line integration: change the base URL, or send an async log through the OpenLLMetry instrumentation.",[142],"Cost tracking computed from token usage and a public price database covering more than 300 models and providers.",[144],"Gateway features on the request path: edge cache, custom rate limits by request count, cost or property, and automatic fallbacks.",[146],"Operational surface: sessions, per-user metrics, custom properties, scores, datasets, webhooks and an MCP server over the data.",[148],"Self-hosting that runs five services: a web frontend, a Cloudflare Worker, the Jawn collector, Supabase and ClickHouse, plus MinIO for bodies.",{"type":128,"level":129,"id":93,"text":94},{"type":121,"content":151},[152],"The request path is deliberately thin. Unless a header enables a feature, the worker forwards the request and returns the response untouched; after the response is complete, the proxy ships logs to Kafka for a separate service to consume. The availability documentation states the design intent directly: all business logic falls back to plain proxying on any error, so a bug in observability degrades to a working relay.",{"type":154,"attrs":155,"inner":159,"caption":160},"diagram",{"viewBox":156,"role":157,"aria-labelledby":158},"0 0 720 372","img","hel-t hel-d","\u003Ctitle id=\"hel-t\">The Helicone request path\u003C\u002Ftitle>\u003Cdesc id=\"hel-d\">Top row: the application sends a chat completion through the OpenAI SDK to the Helicone gateway, which sits on Cloudflare Workers and either serves a cache hit or forwards to the provider and returns the response. Bottom row: after the response is complete, logs go to Kafka and then into ClickHouse for analytics and MinIO for request and response bodies, where the dashboard, HQL, alerts and reports read from. Below them a bar: the async path bypasses the gateway entirely, logging after the response from the application itself, and gives up caching, rate limits and retries in exchange.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">The Helicone request path\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">proxy first, logs second\u003C\u002Ftext>\u003Crect x=\"20\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"102\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Your app\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">OpenAI SDK, one line\u003C\u002Ftext>\u003Cpath d=\"M184 78 H216\" class=\"d-line\" \u002F>\u003Ctext x=\"200\" y=\"70\" text-anchor=\"middle\" class=\"d-small\">request\u003C\u002Ftext>\u003Crect x=\"216\" y=\"46\" width=\"184\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"308\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Gateway\u003C\u002Ftext>\u003Ctext x=\"308\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">Cloudflare Workers, edge\u003C\u002Ftext>\u003Cpath d=\"M400 78 H432\" class=\"d-line\" \u002F>\u003Ctext x=\"416\" y=\"70\" text-anchor=\"middle\" class=\"d-small\">forward\u003C\u002Ftext>\u003Crect x=\"432\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"514\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Provider\u003C\u002Ftext>\u003Ctext x=\"514\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">100+ models, one API\u003C\u002Ftext>\u003Cpath d=\"M596 78 H628\" class=\"d-line\" \u002F>\u003Ctext x=\"612\" y=\"70\" text-anchor=\"middle\" class=\"d-small\">response\u003C\u002Ftext>\u003Crect x=\"216\" y=\"150\" width=\"184\" height=\"64\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"308\" y=\"177\" text-anchor=\"middle\" class=\"d-small\">after the response only\u003C\u002Ftext>\u003Ctext x=\"308\" y=\"199\" text-anchor=\"middle\" class=\"d-small\">logs to Kafka\u003C\u002Ftext>\u003Cpath d=\"M308 110 V150\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"150\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"102\" y=\"177\" text-anchor=\"middle\" class=\"d-small\">edge cache\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"199\" text-anchor=\"middle\" class=\"d-small\">rate limits, fallbacks\u003C\u002Ftext>\u003Cpath d=\"M184 182 H216\" class=\"d-line\" \u002F>\u003Crect x=\"432\" y=\"150\" width=\"268\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"566\" y=\"177\" text-anchor=\"middle\" class=\"d-small\">ClickHouse for analytics\u003C\u002Ftext>\u003Ctext x=\"566\" y=\"199\" text-anchor=\"middle\" class=\"d-small\">MinIO for request bodies\u003C\u002Ftext>\u003Cpath d=\"M400 182 H432\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"254\" width=\"680\" height=\"64\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"360\" y=\"281\" text-anchor=\"middle\" class=\"d-text\">Dashboard, HQL, alerts, reports, MCP\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"303\" text-anchor=\"middle\" class=\"d-small\">Async path skips the gateway: logs after the response, loses caching, rate limits and retries\u003C\u002Ftext>\u003Crect x=\"20\" y=\"336\" width=\"680\" height=\"28\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"360\" y=\"355\" text-anchor=\"middle\" class=\"d-small\">On error the worker falls back to plain proxying, so observability degrades to a relay\u003C\u002Ftext>",[161],"The gateway is on the critical path only in the proxy mode, and the logging happens after the response in both.",{"type":121,"content":163},[164],"One consequence is worth stating plainly: in proxy mode every prompt and every completion transits a third-party edge network by default. The self-hosted deployment removes that, but it also removes the part of the product that makes Helicone easy, because a self-hosted Jawn no longer proxies at all — the docs record that the gateway routes were removed and the AI Gateway must be deployed separately.",{"type":128,"level":129,"id":96,"text":97},{"type":121,"content":167},[168],"The whole integration is a base URL. The example below also switches on the edge cache, which is the fastest way to see the platform do something visible: the second identical request is served from Cloudflare's KV store rather than from the provider.",{"type":170,"code":171},"code","import os\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"https:\u002F\u002Fai-gateway.helicone.ai\",\n    api_key=os.environ[\"HELICONE_API_KEY\"],\n    default_headers={\n        \"Helicone-Cache-Enabled\": \"true\",\n        \"Cache-Control\": \"max-age=3600\",\n        \"Helicone-Cache-Seed\": \"user-123\",\n        \"Helicone-Property-Env\": \"production\",\n    },\n)\n\nresponse = client.chat.completions.create(\n    model=\"gpt-4o-mini\",\n    messages=[{\"role\": \"user\", \"content\": \"Summarise the refund policy.\"}],\n)\n\nprint(response.choices[0].message.content)\n# The next identical call returns Helicone-Cache: HIT and skips the provider.",{"type":121,"content":173},[174],"The cache key is a hash of the cache seed, the request URL, the whole request body and the relevant headers, which makes it exact-match rather than semantic: change one word or one temperature and it misses. Durations run from an hour to a 365-day maximum, buckets up to 20 stored responses are available for non-deterministic prompts, and Helicone-Cache-Ignore-Keys excludes JSON fields such as a request id or a provider prompt_cache_key from the key.",{"type":176,"variant":177,"title":178,"body":179},"callout","warn","Cache keys leak context",[180],[181],"Because the key includes the request body, a per-user cache seed still derives keys from that user's prompt content, and the responses live in Cloudflare Workers KV rather than in infrastructure you control. For anything containing personal data, leave the cache off and decide deliberately whether Helicone belongs in the request path at all.",{"type":128,"level":129,"id":99,"text":100},{"type":121,"content":184},[185],"This is the decision that shapes everything else. Helicone documents both modes and is unusually honest about the gap between them, which makes the choice easy to reason about even though the answer is rarely comfortable.",{"type":187,"head":188,"rows":195},"table",[189,191,193],[190],"Capability",[192],"Proxy mode",[194],"Async mode",[196,203,209,214,219,224,231,236,241,248],[197,199,201],[198],"On the critical path",[200],"yes",[202],"no",[204,206,207],[205],"Gateway cache and buckets",[200],[208],"not available",[210,212,213],[211],"Custom rate limits",[200],[208],[215,217,218],[216],"Automatic retries",[200],[208],[220,222,223],[221],"Prompt auto-formatting",[200],[208],[225,227,229],[226],"Setup effort",[228],"one base URL",[230],"instrument with OpenLLMetry",[232,234,235],[233],"Sessions and user metrics",[200],[200],[237,239,240],[238],"Custom properties and scores",[200],[200],[242,244,246],[243],"Data path",[245],"prompts cross Helicone's edge",[247],"prompts stay with the provider",[249,251,253],[250],"Use it when",[252],"you want caching and rate limits",[254],"propagation delay is unacceptable",{"type":121,"content":256},[257],"The reasonable conclusion is that these are two products sharing a name. The proxy is a gateway with analytics attached; the async integration is a logging library with a web UI. Teams that take the proxy get caching, rate limiting and fallbacks that they would otherwise build, and pay for it with an extra hop and a data-residency decision. Teams that take the async path keep their latency budget and their network boundary, and build the gateway themselves or go without it.",{"type":128,"level":129,"id":102,"text":103},{"type":121,"content":260},[261],"The published benchmark is small but unusually well specified: 500 requests with unique prompts interleaved between OpenAI and Helicone inside the same one-second window, alternating which endpoint was called first, with the prompt context maximised, on text-ada-001, logging round-trip latency for both sets.",{"type":187,"head":263,"rows":270},[264,266,268],[265],"Statistic",[267],"OpenAI direct (s)",[269],"Helicone proxied (s)",[271,277,284,290,297,304,311,318,325],[272,274,276],[273],"Mean",[275],"2.21",[275],[278,280,282],[279],"Median",[281],"2.87",[283],"2.90",[285,287,289],[286],"Standard deviation",[288],"1.12",[288],[291,293,295],[292],"p90",[294],"3.27",[296],"3.29",[298,300,302],[299],"Maximum",[301],"3.56",[303],"3.76",[305,307,309],[306],"Method",[308],"500 interleaved requests",[310],"same prompts, alternating order",[312,314,316],[313],"Model",[315],"text-ada-001",[317],"text-ada-001, max context",[319,321,323],[320],"Overhead",[322],"baseline",[324],"visible only in the maximum",[326,328,330],[327],"Who ran it",[329],"vendor",[329],{"type":121,"content":332},[333],"Read the maximum row carefully. A 200 millisecond difference on a three-and-a-half second response is 6 per cent, and the same absolute hop applied to a 300 millisecond model call or to a 40 millisecond tool call would dominate it. The benchmark is also a vendor test on a retired model, so it demonstrates that the overhead is bounded, not what it will be for a given workload.",{"type":176,"variant":335,"title":336,"body":337},"note","Where the time actually goes",[338],[339],"Logging is not the cost. The proxy writes to Kafka only after the response is returned, so telemetry does not sit in front of the user. The cost is the round trip to an edge location that is close to the caller but not necessarily close to the provider's region, plus the gateway features you switch on, which by definition do work per request.",{"type":128,"level":129,"id":105,"text":106},{"type":121,"content":342},[343],"The free Hobby plan covers 10,000 requests a month, one seat, one organisation, one gigabyte and seven days of retention. Pro is $79 a month for unlimited seats, one organisation, alerts, reports, HQL, one month of retention and 1,000 ingested logs per minute. Team is $799 for five organisations, SOC 2 and HIPAA options, three months of retention and 15,000 logs per minute. Enterprise adds unlimited organisations, SAML SSO and on-prem deployment.",{"type":121,"content":345},[346],"The number to watch is the one nobody prices prominently: ingestion. A Hobby workspace ingests 10 logs a minute, Pro 1,000, Team 15,000. A production application can emit several logs per user request — a model call, a retrieval step, a tool call — so a moderately busy product exhausts the Pro ceiling before it exhausts the request allowance, and the plan that fixes it costs 799 dollars a month. Storage is metered separately above the first gigabyte.",{"type":187,"head":348,"rows":377},[349,355,361,367,372],[350,351,352,353,354],"Plan","Price","Requests","Ingestion","Retention",[356,357,358,359,360],"Hobby","$0","10,000 per month","10 logs per minute","7 days",[362,363,364,365,366],"Pro","$79 per month","usage-based","1,000 logs per minute","1 month",[368,369,364,370,371],"Team","$799 per month","15,000 logs per minute","3 months",[373,374,364,375,376],"Enterprise","custom","30,000 logs per minute","forever",[378,387,395,406],[379,381,383,385,386],[380],"Seats",[382],"1",[384],"unlimited",[384],[384],[388,390,391,392,394],[389],"Organisations",[382],[382],[393],"5",[384],[396,398,400,402,404],[397],"Notable additions",[399],"nothing",[401],"HQL, alerts, reports",[403],"SOC 2, HIPAA",[405],"SAML SSO, on-prem",[407,409,411,412,413],[408],"Self-host",[410],"free",[410],[410],[414],"Helm chart on request",{"type":128,"level":129,"id":108,"text":109},{"type":121,"content":417},[418],"The Apache-2.0 repository contains the entire platform, and the README is direct about the operational shape: five services, and a manual deployment that it marks as not recommended in favour of the Docker compose file or an Enterprise Helm chart.",{"type":134,"ordered":135,"items":420},[421,423,425,427,429,431],[422],"Web: the dashboard frontend, a Next.js application.",[424],"Worker: the proxy, deployed as Cloudflare Workers in the hosted version.",[426],"Jawn: the log collector, an Express service with Tsoa-generated routes.",[428],"Supabase: the application database and authentication.",[430],"ClickHouse: the analytics store that answers cost and latency questions.",[432],"MinIO: object storage for request and response bodies.",{"type":176,"variant":177,"title":434,"body":435},"Two things the docs get right to warn about",[436],[437],"Jawn no longer proxies: the gateway routes were removed, so a self-hosted deployment runs the AI Gateway as a separate service or does not proxy at all. And port 8585 accepts proxy requests without authentication, so anyone who can reach it can spend your provider credits — restrict it at the firewall, put TLS in front, and mount volumes for Postgres, ClickHouse and MinIO or a restart wipes the data.",{"type":128,"level":129,"id":111,"text":112},{"type":121,"content":440},[441],"The honest weaknesses are structural, not cosmetic. Helicone's data model is built around requests that pass through its gateway, so a team that already traces with OpenTelemetry spans has to choose between two instrumentation stacks. The evaluation story is thin next to Langfuse or Braintrust: prompts, playground and scores exist, but there is no first-class experiments workflow. And the ownership question is live — Helicone joined Mintlify in 2025, which removes the open-core lock-in risk but adds the usual vendor dependency.",{"type":187,"head":443,"rows":452},[444,446,448,450],[445],"",[447],"Helicone",[449],"Langfuse",[451],"Arize Phoenix",[453,462,471,480,489,496,504,512],[454,456,458,460],[455],"Licence",[457],"Apache 2.0",[459],"MIT core, enterprise extras",[461],"Elastic License 2.0",[463,465,467,469],[464],"Instrumentation",[466],"proxy or OpenLLMetry",[468],"OpenTelemetry native",[470],"OpenInference on OTel",[472,474,476,478],[473],"Strongest suit",[475],"one-line start, gateway features",[477],"evaluation and prompt workflows",[479],"local notebook-style evaluation",[481,483,485,487],[482],"Weakest suit",[484],"not OTel-native for tracing",[486],"heavier to start than a base URL",[488],"ELv2 is not permissive",[490,492,494,495],[491],"Gateway features",[493],"cache, limits, fallbacks",[202],[202],[497,498,500,502],[408],[499],"Apache 2.0, five services",[501],"MIT, Docker Compose or Helm",[503],"ELv2, self-hostable",[505,506,508,510],[11],[507],"built in",[509],"configurable per model",[511],"separate product",[513,515,517,519],[514],"Best fit",[516],"teams wanting visibility this week",[518],"teams running evaluations",[520],"teams tracing and evaluating locally",{"type":121,"content":522},[523],"Langfuse is MIT-licensed, was acquired by ClickHouse in January 2026 with the licence and self-hosting explicitly unchanged, and stores traces in ClickHouse as well; if the deciding factor is a permissive licence with a real evaluation stack, it is the stronger tool. Arize Phoenix is built on OpenInference conventions and runs anywhere, including air-gapped, but its Elastic License 2.0 is not a permissive one. Helicone's argument is not capability, it is time: the distance between a repository and a cost dashboard measured in requests per user is one line of code.",{"type":128,"level":129,"id":114,"text":115},{"type":121,"content":526},[527],"Helicone is the right tool when nobody can answer what the LLM spend was last Tuesday, and the wrong tool when latency, data residency or tracing standards decide the architecture. Its engineering is unremarkable in the best sense: the proxy is thin, logs go out after the response, and the product is Apache-2.0 from the gateway to the dashboard. Its weakness is equally plain: it wants to be your gateway, and everything that makes it comfortable is a consequence of that.",{"type":134,"ordered":529,"items":530},true,[531,533,535,537,539,541],[532],"Use it when the team needs per-request cost, latency and error visibility within a week and has no tracing stack yet.",[534],"Use it when the edge cache, custom rate limits and automatic fallbacks are worth having on the request path.",[536],"Use the async OpenLLMetry path when you want Helicone's analytics without a proxy hop, and accept losing the gateway features.",[538],"Avoid it when prompts must not leave your network or when the proxy hop eats a hard latency budget.",[540],"Avoid it when the platform already traces with OpenTelemetry and a second instrumentation model would fragment the data.",[542],"Reconsider it above roughly 1,000 ingested logs a minute, or when the request allowance is large enough that the usage meter needs its own line in the budget.",{"type":176,"variant":544,"title":545,"body":546},"tip","The one thing to measure",[547],[548],"Before committing, measure the gateway hop against your own latency budget with your own model and region. Helicone publishes a fair benchmark, but a fair benchmark on text-ada-001 tells you almost nothing about a 300 millisecond streaming response served from a region three time zones away from the provider.",{"type":128,"level":129,"id":117,"text":118},{"type":134,"ordered":529,"items":551},[552,556,559,562,565,568,571,574,577],[553],{"tag":554,"href":31,"children":555},"a",[30],[557],{"tag":554,"href":34,"children":558},[33],[560],{"tag":554,"href":37,"children":561},[36],[563],{"tag":554,"href":40,"children":564},[39],[566],{"tag":554,"href":43,"children":567},[42],[569],{"tag":554,"href":46,"children":570},[45],[572],{"tag":554,"href":49,"children":573},[48],[575],{"tag":554,"href":52,"children":576},[51],[578],{"tag":554,"href":55,"children":579},[54],[581,630,713,770],{"slug":582,"published":583,"minutes":584,"category":7,"tags":585,"keywords":591,"about":598,"sources":602,"cover":621,"og":622,"expertise":58,"locales":623,"lang":60,"title":624,"description":625,"coverAlt":626,"url":627,"pricing":628,"kind":629},"ollama","2026-09-29",11,[586,587,588,589,590],"Local inference","Open models","llama.cpp","GGUF","Model serving",[582,592,593,594,595,596,597],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[599],{"name":600,"url":601},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[603,606,609,612,615,618],{"title":604,"url":605},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":607,"url":608},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":610,"url":611},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":613,"url":614},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":616,"url":617},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":619,"url":620},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[60,61,62],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":631,"published":632,"minutes":584,"category":7,"tags":633,"keywords":639,"about":647,"sources":654,"cover":706,"og":707,"expertise":58,"locales":708,"lang":60,"title":709,"description":710,"coverAlt":711,"url":650,"pricing":712,"kind":634},"portkey","2026-09-28",[634,635,636,637,638],"LLM gateway","Guardrails","Routing","Observability","Cost control",[640,641,642,643,644,645,646],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[648,651],{"name":649,"url":650},"Portkey","https:\u002F\u002Fportkey.ai",{"name":652,"url":653},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[655,658,661,664,667,670,673,676,679,682,685,688,691,694,697,700,703],{"title":656,"url":657},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":659,"url":660},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":662,"url":663},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":665,"url":666},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":668,"url":669},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":671,"url":672},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":674,"url":675},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":677,"url":678},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":680,"url":681},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":683,"url":684},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":686,"url":687},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":689,"url":690},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":692,"url":693},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":695,"url":696},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":698,"url":699},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":701,"url":702},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":704,"url":705},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[60,61,62],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":714,"published":715,"minutes":6,"category":7,"tags":716,"keywords":720,"about":728,"sources":738,"cover":763,"og":764,"expertise":58,"locales":765,"lang":60,"title":766,"description":767,"coverAlt":768,"url":730,"pricing":769,"kind":9},"langfuse","2026-08-13",[9,717,10,718,719],"Tracing","Self-hosting","Evaluation",[714,721,722,723,724,725,726,727],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[729,731,732,735],{"name":449,"url":730},"https:\u002F\u002Flangfuse.com",{"name":10,"url":24},{"name":733,"url":734},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":736,"url":737},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[739,742,745,748,751,754,757,760],{"title":740,"url":741},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":743,"url":744},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":746,"url":747},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":749,"url":750},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":752,"url":753},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":755,"url":756},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":758,"url":759},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":761,"url":762},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[60,61,62],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":771,"published":772,"minutes":6,"category":7,"tags":773,"keywords":778,"about":785,"sources":789,"cover":805,"og":806,"expertise":58,"locales":807,"lang":60,"title":808,"description":809,"coverAlt":810,"url":811,"pricing":812,"kind":634},"openrouter","2026-07-23",[634,774,775,776,777],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[771,779,642,780,781,782,783,784],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[786],{"name":787,"url":788},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[790,793,794,797,800,803],{"title":791,"url":792},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":704,"url":705},{"title":795,"url":796},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":798,"url":799},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":801,"url":802},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":804,"url":788},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[60,61,62],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",1791383548969]