[{"data":1,"prerenderedAt":699},["ShallowReactive",2],{"tool-portkey-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":29,"cover":81,"og":82,"expertise":83,"locales":84,"lang":85,"title":88,"description":89,"coverAlt":90,"url":25,"pricing":91,"kind":9,"metaTitle":92,"takeaways":93,"faq":99,"toc":112,"blocks":136,"others":492},"portkey","2026-09-28",11,"llmops",[9,10,11,12,13],"LLM gateway","Guardrails","Routing","Observability","Cost control",[15,16,17,18,19,20,21],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[23,26],{"name":24,"url":25},"Portkey","https:\u002F\u002Fportkey.ai",{"name":27,"url":28},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[30,33,36,39,42,45,48,51,54,57,60,63,66,69,72,75,78],{"title":31,"url":32},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":34,"url":35},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":37,"url":38},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":40,"url":41},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":43,"url":44},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":46,"url":47},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":49,"url":50},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":52,"url":53},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":55,"url":56},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":58,"url":59},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":61,"url":62},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":64,"url":65},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":67,"url":68},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":70,"url":71},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":73,"url":74},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":76,"url":77},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":79,"url":80},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg","ai-engineer",[85,86,87],"en","de","hu","Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month","Portkey review: routing, guardrails and latency · Balázs Csorba",[94,95,96,97,98],"Portkey's routing config is the strongest part of the product: a closed schema, four nestable strategy modes, weights that normalise to 100 and a circuit breaker whose cooldown cannot be set below 30 seconds.","The hosted gateway is a network hop. Portkey's own benchmark repository measures it at plus 93 ms on average against a direct Bedrock call, and calls 50 to 150 ms typical for the extra two hops.","Guardrails only ever read the last message and never read image inputs, and a denied synchronous check returns 446, a status code most SDKs will retry as if it were a transport error.","Pricing is per recorded log rather than per request, and the pricing page and the caching documentation disagree about whether the 49 dollar Production plan includes semantic caching.","Since the May 2026 acquisition it ships as the gateway inside Prisma AIRS, which changes who signs the contract without changing the API.",[100,103,106,109],{"q":101,"a":102},"What is Portkey and what does an LLM gateway actually do?","An LLM gateway is a proxy that sits between your application and the model providers it calls. Portkey's turns provider keys, retries, fallbacks, weighted routing, caching, guardrails and cost attribution into JSON configuration objects attached to a request, so an OpenAI-compatible client keeps working while the routing policy changes in a dashboard instead of a deploy.",{"q":104,"a":105},"Does Portkey add latency to model calls?","On the hosted product, yes. Portkey's own benchmark repository measures an average of plus 93 ms and a median of plus 25 ms routing through the cloud gateway against a direct Bedrock call, and its README describes 50 to 150 ms as typical for the extra two hops. Self-hosted it is a different number: the gateway repository claims under 1 ms of processing.",{"q":107,"a":108},"Is Portkey open source?","The gateway is MIT licensed and runs from a single npm command, and the published package currently sits at version 1.15.2 with a 2.0 pre-release branch in progress. The control plane, the dashboard, the Model Catalog and the enterprise deployment options are commercial. Running the proxy yourself does not give you the managed product.",{"q":110,"a":111},"Portkey or LiteLLM?","LiteLLM is MIT licensed with no managed tier, so it wins on cost and on data staying inside your own network, and you build the dashboards yourself. Portkey wins when you want routing policy, guardrails and stored credentials to live in a product someone else operates, and you accept 49 dollars a month for 100,000 recorded logs plus 9 dollars per additional 100,000, along with a network hop.",[113,116,119,122,124,127,130,133],{"id":114,"title":115},"what-it-is","What it is",{"id":117,"title":118},"how-it-works","How it works",{"id":120,"title":121},"getting-started","Getting started",{"id":123,"title":10},"guardrails",{"id":125,"title":126},"pricing-and-latency","Pricing and latency",{"id":128,"title":129},"where-it-shingles","Where it does not fit",{"id":131,"title":132},"verdict","Verdict",{"id":134,"title":135},"sources","Sources",[137,141,144,147,150,189,190,193,202,205,208,210,217,218,221,223,230,237,238,241,259,265,266,269,307,310,313,351,356,359,360,363,409,412,413,416,431,437,438],{"type":138,"content":139},"paragraph",[140],"Portkey is an LLM gateway: a proxy that sits between an application and the model providers it calls, and turns provider keys, retries, fallbacks, caching, guardrails and cost attribution into configuration objects rather than application code. It is one of the more complete gateways on the market, and since Palo Alto Networks closed its acquisition of Portkey in May 2026 it also ships as the gateway inside Prisma AIRS. The position taken here: the routing layer is the strongest in its class and worth the operational dependency, while the hosted deployment and the guardrail semantics are where the caveats sit.",{"type":138,"content":142},[143],"What it replaces is the hand-rolled retry wrapper every team writes in the first month of an LLM product, which is usually a try block with a sleep, one hard-coded provider key and no record of what any of it cost. Portkey moves that behind a single OpenAI-compatible base URL, so the OpenAI SDK, the Anthropic SDK, LangChain or a raw fetch call all talk to the same endpoint. It competes directly with LiteLLM, free and self-hosted; with OpenRouter, a paid marketplace rather than a control plane; and with Cloudflare AI Gateway, which is close to free because it rides infrastructure most teams already pay for.",{"type":145,"level":146,"id":114,"text":115},"heading",2,{"type":138,"content":148},[149],"Portkey is really two products sharing one repository. The gateway itself is an MIT-licensed Node proxy that runs from a single command and listens on port 8787 with a local console attached; the control plane around it is a paid service that stores credentials, hosts the configuration UI and keeps the logs. That split explains most of what follows. What the proxy does is well documented and free, and what the control plane does is where the pricing, the guardrail catalogue and the compliance story live.",{"type":151,"ordered":152,"items":153},"list",false,[154,169,171,177,183,185,187],[155,156,160,161,164,165,168],"Runs as one command, ",{"tag":157,"children":158},"code",[159],"npx @portkey-ai\u002Fgateway",", serving on ",{"tag":157,"children":162},[163],"localhost:8787"," with a console at ",{"tag":157,"children":166},[167],"\u002Fpublic\u002F",".",[170],"MIT licensed with roughly 13,100 GitHub stars and 1,300 forks; the published npm package sits at version 1.15.2 and a 2.0 pre-release branch is in progress.",[172,173,176],"One OpenAI-compatible endpoint, plus Anthropic's ",{"tag":157,"children":174},[175],"\u002Fv1\u002Fmessages"," and the Open Responses format, so the model string is the only thing that changes when the provider does.",[178,179,182],"Model strings carry the provider: ",{"tag":157,"children":180},[181],"@openai-prod\u002Fgpt-4o"," resolves to a stored integration, along with its budget and its rate limit.",[184],"Strategy objects nest. A fallback can contain a load balancer that contains another fallback, each with its own weights, status-code triggers and circuit breaker.",[186],"Guardrails evaluate inputs and outputs, running either asynchronously with no added latency or synchronously with documented deny codes of 246 and 446.",[188],"Owned by Palo Alto Networks since May 2026 and sold as the Prisma AIRS AI Gateway, generally available since 16 July 2026.",{"type":145,"level":146,"id":117,"text":118},{"type":138,"content":191},[192],"A request arrives carrying a Portkey API key and, usually, a config. The config names a strategy and a list of targets. The gateway resolves each target to a stored provider integration, applies caches and guardrails as the config dictates, forwards to the chosen provider, and writes a log line with latency, token counts and cost. Where this differs from a hand-rolled proxy is not the forwarding. It is that the decision is data: the same JSON runs unchanged in the hosted product, in the open-source proxy and in a self-hosted data plane.",{"type":194,"attrs":195,"inner":199,"caption":200},"diagram",{"viewBox":196,"role":197,"aria-labelledby":198},"0 0 620 300","img","pk-flow-t pk-flow-d","\u003Ctitle id=\"pk-flow-t\">A request through the Portkey gateway\u003C\u002Ftitle>\u003Cdesc id=\"pk-flow-d\">An application sends an OpenAI-compatible request to the Portkey gateway. The gateway applies the attached config, then fans the request out to one of three providers using retry, fallback and weighted routing. A guardrail verdict is checked synchronously and returns 246 or 446. Every call writes a log line with latency, tokens and cost.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-pk\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003Cmarker id=\"ah-pk-a\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head-accent\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Crect x=\"20\" y=\"100\" width=\"130\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"85\" y=\"128\" text-anchor=\"middle\" class=\"d-text\">your app\u003C\u002Ftext>\u003Ctext x=\"85\" y=\"150\" text-anchor=\"middle\" class=\"d-small\">OpenAI SDK\u003C\u002Ftext>\u003Crect x=\"200\" y=\"80\" width=\"170\" height=\"104\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"285\" y=\"116\" text-anchor=\"middle\" class=\"d-text\">Portkey gateway\u003C\u002Ftext>\u003Ctext x=\"285\" y=\"140\" text-anchor=\"middle\" class=\"d-small\">config, retry, cache\u003C\u002Ftext>\u003Ctext x=\"285\" y=\"160\" text-anchor=\"middle\" class=\"d-small\">guardrail, log\u003C\u002Ftext>\u003Crect x=\"20\" y=\"230\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"95\" y=\"258\" text-anchor=\"middle\" class=\"d-text\">guardrail\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"280\" text-anchor=\"middle\" class=\"d-small\">246, or 446\u003C\u002Ftext>\u003Crect x=\"200\" y=\"230\" width=\"170\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"285\" y=\"258\" text-anchor=\"middle\" class=\"d-text\">logs\u003C\u002Ftext>\u003Ctext x=\"285\" y=\"280\" text-anchor=\"middle\" class=\"d-small\">cost, latency\u003C\u002Ftext>\u003Crect x=\"450\" y=\"40\" width=\"140\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"520\" y=\"74\" text-anchor=\"middle\" class=\"d-text\">OpenAI\u003C\u002Ftext>\u003Crect x=\"450\" y=\"126\" width=\"140\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"520\" y=\"160\" text-anchor=\"middle\" class=\"d-text\">Anthropic\u003C\u002Ftext>\u003Crect x=\"450\" y=\"212\" width=\"140\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"520\" y=\"246\" text-anchor=\"middle\" class=\"d-text\">Bedrock\u003C\u002Ftext>\u003Cpath d=\"M150 132H198\" class=\"d-line-accent\" marker-end=\"url(#ah-pk-a)\" \u002F>\u003Cpath d=\"M370 104H400V68H448\" class=\"d-line\" marker-end=\"url(#ah-pk)\" \u002F>\u003Cpath d=\"M370 132H448\" class=\"d-line\" marker-end=\"url(#ah-pk)\" \u002F>\u003Cpath d=\"M370 160H400V240H448\" class=\"d-line\" marker-end=\"url(#ah-pk)\" \u002F>\u003Ctext x=\"402\" y=\"120\" text-anchor=\"middle\" class=\"d-label\">retry, fallback\u003C\u002Ftext>\u003Cpath d=\"M250 184V210H95V228\" class=\"d-line d-dash\" marker-end=\"url(#ah-pk)\" \u002F>\u003Ctext x=\"95\" y=\"204\" text-anchor=\"middle\" class=\"d-label\">sync\u003C\u002Ftext>\u003Cpath d=\"M320 184V228\" class=\"d-line d-dash\" marker-end=\"url(#ah-pk)\" \u002F>\u003Ctext x=\"376\" y=\"204\" text-anchor=\"middle\" class=\"d-label\">every call\u003C\u002Ftext>",[201],"One request through the gateway: the config decides the target, a synchronous guardrail can deny with 246 or 446, and every call is logged with its latency, tokens and cost.",{"type":138,"content":203},[204],"The interesting branch is the guardrail. Run asynchronously, which is the default, the check runs alongside the model call, the result is only logged, and the provider's own status code comes back unchanged. Run synchronously, the check blocks: a pass returns 200, a failure returns 246 when deny is off and 446 when it is on. Both codes sit outside the range any client library was written against. A library that treats any status other than 200 as success will silently accept a 246 that was supposed to be flagged, and one that treats anything outside the 2xx range as an exception will throw on a 446 and retry it as though the network had failed. That is the sharpest edge in the product.",{"type":138,"content":206},[207],"The config object is the other half. Its schema is closed, with a fixed enum of four strategy modes, single, loadbalance, fallback and conditional, and a fixed set of keys, so a typo is rejected rather than silently ignored. Targets are themselves configs, which is what makes the strategies compose.",{"type":157,"code":209},"{\n  \"strategy\": { \"mode\": \"fallback\", \"on_status_codes\": [429, 500, 503] },\n  \"retry\": { \"attempts\": 3, \"use_retry_after_headers\": true },\n  \"cb_config\": { \"failure_threshold\": 5, \"cooldown_interval\": 60000 },\n  \"targets\": [\n    { \"provider\": \"@openai-prod\",\n      \"override_params\": { \"model\": \"gpt-4o\" } },\n    { \"strategy\": { \"mode\": \"loadbalance\" },\n      \"targets\": [\n        { \"provider\": \"@anthropic-prod\", \"weight\": 0.8,\n          \"override_params\": { \"model\": \"claude-sonnet-4-5-20250929\" } },\n        { \"provider\": \"@bedrock-prod\", \"weight\": 0.2,\n          \"override_params\": { \"model\": \"anthropic.claude-3-5-sonnet-20241022-v2:0\" } }\n      ] }\n  ]\n}",{"type":138,"content":211},[212,213,216],"Three details earn their own attention. Weights are normalised to 100 and a weight of zero keeps a target in the config while sending it no traffic, which is how a canary is paused rather than deleted. The circuit breaker's ",{"tag":157,"children":214},[215],"cooldown_interval"," has a floor of 30,000 milliseconds, so a fast-flap protection loop cannot be configured into a tight retry storm. And sticky routing hashes the fields you name with a one-hour default TTL, but its two-tier cache is in-memory plus Redis, so without Redis it works on a single instance only.",{"type":145,"level":146,"id":120,"text":121},{"type":138,"content":219},[220],"Install the SDK, add a provider in the Model Catalog, and change one string. The Portkey SDK is a superset of the OpenAI client, so an existing integration usually needs nothing but its base URL and key swapped.",{"type":157,"code":222},"from portkey_ai import Portkey\n\nclient = Portkey(api_key=\"PORTKEY_API_KEY\")\n\nanswer = client.chat.completions.create(\n    model=\"@openai-prod\u002Fgpt-4o\",              # @provider-slug\u002Fmodel-name\n    messages=[{\"role\": \"user\", \"content\": \"Summarise this ticket in one line.\"}],\n)\nprint(answer.choices[0].message.content)\n\n# Same code, different model: only the string changes\nanswer = client.chat.completions.create(\n    model=\"@anthropic-prod\u002Fclaude-sonnet-4-5-20250929\",\n    max_tokens=512,\n    messages=[{\"role\": \"user\", \"content\": \"Summarise this ticket in one line.\"}],\n)\nprint(answer.choices[0].message.content)",{"type":138,"content":224},[225,226,229],"Two things matter before this goes near production. The provider slug is resolved server-side from the Model Catalog rather than from a literal in the request, so a wrong model string fails at request time and not at deploy time. And the config that carries retries, caching and guardrails is not in the snippet: it is attached either as a config ID on the client or as a JSON blob in the ",{"tag":157,"children":227},[228],"x-portkey-config"," header, which is exactly what lets routing policy change without a deploy.",{"type":231,"variant":232,"title":233,"body":234},"callout","warn","Config changes are production changes",[235],[236],"A config edited in the dashboard is production behaviour with no pull request behind it. Version the JSON, diff it before every change and keep a known-good copy, because a malformed fallback chain fails open or fails closed depending on which field is wrong.",{"type":145,"level":146,"id":123,"text":10},{"type":138,"content":239},[240],"Guardrails are checks attached to a request and evaluated on the input, the output or both. They are the part of Portkey most likely to be misconfigured, partly because the feature list reads wider than the behaviour. The documentation is unusually explicit about the limits, which at least makes the limits easy to find.",{"type":151,"ordered":152,"items":242},[243,245,247,249,255,257],[244],"Only the last message in the request is evaluated, and only its text portions. Image inputs, base64 or URL, are not checked at all.",[246],"Guardrails do not run on the Assistants, Audio, Images, Files, Batch, Fine-tuning, Moderations or Models endpoints. They do run on chat completions, completions, embeddings with input only, messages, responses and prompt completions.",[248],"Output guardrails on a streamed response are informational: the verdict arrives as a trailing chunk after the done marker and triggers no fallback and no retry.",[250,251,254],"To see hook results in a stream at all, strict OpenAI compliance has to be switched off with the ",{"tag":157,"children":252},[253],"x-portkey-strict-open-ai-compliance"," header, because the default strips them.",[256],"Tiers follow the plan: basic checks on Developer, basic plus partner and pro checks on Production, everything including custom on Enterprise. Most checks are deterministic, regex, JSON schema, word and character counts, with LLM-based checks such as prompt-injection scanning on top.",[258],"Partner guardrails exist, including Aporia, Pillar Security, SydeLabs, Zscaler AI Guard and Akto, each called over HTTP with its own timeout: 10,000 ms for Zscaler, 5,000 ms for Akto.",{"type":138,"content":260},[261,262,264],"The Anthropic path carries one more wrinkle. On ",{"tag":157,"children":263},[175]," the hook results arrive as a dedicated event the Anthropic SDK does not parse, so reading them means dropping down to cURL. That is the kind of small inconsistency that decides whether a feature gets adopted or quietly ignored, and it is worth checking against your own SDK before committing to a guardrail-based control.",{"type":145,"level":146,"id":125,"text":126},{"type":138,"content":267},[268],"Latency is where the marketing and the measurement diverge. Three numbers exist for the same product and they are not describing the same thing, which is why the table below names the source of each rather than picking the flattering one.",{"type":270,"head":271,"rows":278},"table",[272,274,276],[273],"Figure",[275],"Source",[277],"What it measures",[279,286,293,300],[280,282,284],[281],"Under 1 ms",[283],"Gateway repository README",[285],"The self-hosted proxy's own processing time, not the hosted round trip",[287,289,291],[288],"Sub-10 ms at 99.9999% uptime",[290],"Vendor blog, October 2025",[292],"A hosting claim covering 10 billion requests a month, with no published method",[294,296,298],[295],"Plus 93 ms average, plus 25 ms median",[297],"Portkey's own benchmark repository",[299],"The cloud gateway against a direct Bedrock call, two workers, three requests per iteration",[301,303,305],[302],"50 to 150 ms typical",[304],"The same repository, overhead section",[306],"What the vendor itself calls the round-trip cost of the two extra network hops",{"type":138,"content":308},[309],"The honest reading is that the two figures describe two different products. Under 1 ms is the open-source proxy's processing time, which is genuinely good and is the main argument for self-hosting it. Sub-10 ms is a hosting claim with no method attached. The plus 93 ms figure is the only one published with a runnable harness, and the same repository describes 50 to 150 ms as typical. Against a 400 ms time to first token that is invisible; against a 90 ms autocomplete or voice path it is the entire budget. Measure it on your own traffic before adopting it, not from the pricing page.",{"type":138,"content":311},[312],"Pricing follows, and it is charged on recorded logs rather than on requests. That is unusual and mostly harmless: the free plan keeps serving traffic after the log cap and simply stops recording. The paid tiers are where it starts to matter, because overage is priced per 100,000 requests on top of a log allowance.",{"type":270,"head":314,"rows":323},[315,317,319,321],[316],"Plan",[318],"Price",[320],"Recorded logs",[322],"Retention and what is included",[324,333,342],[325,327,329,331],[326],"Developer",[328],"Free",[330],"10,000 a month",[332],"3 days for logs, 30 for metrics; 3 prompt templates; deterministic guardrails; community support; requests keep flowing past the cap",[334,336,338,340],[335],"Production",[337],"49 dollars a month",[339],"100,000, then 9 dollars per extra 100,000",[341],"30 days for logs, 90 for metrics; role-based access, service account keys, LLM and partner guardrails, semantic caching, production support",[343,345,347,349],[344],"Enterprise",[346],"Quoted",[348],"10 million or more",[350],"Custom retention; private cloud and VPC hosting, SSO, data export, SOC 2 Type 2, GDPR and HIPAA, data isolation",{"type":231,"variant":232,"title":352,"body":353},"The pricing page and the caching docs disagree",[354],[355],"The plan comparison table lists simple and semantic caching under the 49 dollar Production tier, while the caching documentation states that semantic caching requires a vector database and is available only on select Enterprise plans. Confirm which applies before building a cost model that depends on semantic cache hit rates.",{"type":138,"content":357},[358],"Enterprise is where the product becomes a different thing: a data plane inside your own VPC, Helm charts on Kubernetes 1.20 or later, one to two cores and two to four gigabytes of memory per instance, logs in S3-compatible object storage or MongoDB, and a control plane the gateway syncs from once a minute while holding a seven-day local cache of configs and keys. The documentation recommends a volatile-lru eviction policy so that live configuration survives memory pressure. This is a deployment to operate, not a container to forget about.",{"type":145,"level":146,"id":128,"text":129},{"type":138,"content":361},[362],"The honest weaknesses first, because they rule the tool out for whole classes of team. A hosted gateway is a single point of failure and a single tenancy boundary for every request your application makes, and the status page does show control plane incidents, including two in August 2026 lasting 40 minutes and two hours. Semantic caching is gated behind an Enterprise conversation, so the cost-saving feature teams cite in their business case is not available to them. The guardrail model only ever sees the last message, which means it is not a prompt-injection defence for a multimodal conversation. And the config lives in a SaaS dashboard, which means the routing policy of a system is no longer reviewable in the repository that owns the system.",{"type":270,"head":364,"rows":373},[365,367,369,371],[366],"Tool",[368],"Licence and cost",[370],"Where it runs",[372],"What it does not do",[374,382,391,400],[375,376,378,380],[24],[377],"MIT gateway, paid control plane from 49 dollars",[379],"Vendor cloud, or your own VPC on Enterprise",[381],"Never reads image inputs; semantic cache is Enterprise-only on the hosted plan",[383,385,387,389],[384],"LiteLLM",[386],"MIT, self-host free, Enterprise priced by request capacity",[388],"Your infrastructure only",[390],"No hosted tier, so no managed dashboards, shared guardrails or credential vault",[392,394,396,398],[393],"OpenRouter",[395],"Pay per token, 5.5 percent platform fee on credit purchases",[397],"Their cloud",[399],"No self-hosted gateway; BYOK runs above 25,000 dollars of list-price inference a month",[401,403,405,407],[402],"Cloudflare AI Gateway",[404],"Core features free on every plan, guardrails billed as Workers AI inference",[406],"Cloudflare edge, needs the Workers paid plan at volume",[408],"Guardrail checks are Llama Guard 3 8B on Workers AI, with no bring-your-own",{"type":138,"content":410},[411],"The choice is therefore less about features than about where you want the policy to live. If the routing rules belong in version control next to the code, LiteLLM wins on every axis including cost. If you want a security team to be able to change them, if you want credentials that never touch an application environment, or if you are running inside a regulated environment that already buys from Palo Alto Networks, Portkey's control plane is the argument. Cloudflare AI Gateway is the third option worth pricing: core features are free on every plan and the cost is inference rather than platform, which inverts the calculation whenever traffic is small.",{"type":145,"level":146,"id":131,"text":132},{"type":138,"content":414},[415],"Portkey is the most complete managed LLM gateway available, and the config object alone is a better design than most hand-built equivalents: nestable strategies, a closed schema, a hard floor on the circuit breaker cooldown. Take it when routing policy is a governance problem rather than a coding problem, and accept both the network hop and the fact that your routing config now lives somewhere other than your repository. The opinionated part is that this is the right default for enterprises and the wrong default for a small team, which would be better served by the free tier of the open-source proxy or by LiteLLM and a weekend of wiring.",{"type":151,"ordered":417,"items":418},true,[419,421,423,425,427,429],[420],"Take Portkey when several teams share model credentials and someone has to own budgets, allow-lists and rate limits centrally. That is the product's real job.",[422],"Take it when a security team has to be able to change guardrails and inspect full request logs without a deploy.",[424],"Do not take it for latency-sensitive completion or voice paths until you have measured the hop on your own traffic, because the published overhead is tens of milliseconds and the marketing figure is not.",[426],"Do not rely on guardrails as a prompt-injection control for multimodal traffic. They never read the image, and they only read the last message.",[428],"Skip it if your routing policy has to be reviewed in code review. Keep that config in the repository and run LiteLLM or the open-source gateway instead.",[430],"Reconsider at the point where the control plane becomes the bottleneck: at that volume, a gateway on Cloudflare or an in-house proxy in your own VPC is cheaper and simpler than the enterprise tier.",{"type":231,"variant":432,"title":433,"body":434},"note","The number to check before you buy",[435],[436],"Ask for the overhead figure measured between your region and the provider region you actually use, on a request shape matching your time to first token. Portkey's own repository publishes a harness for exactly this, and it reports a materially worse number than the homepage.",{"type":145,"level":146,"id":134,"text":135},{"type":151,"ordered":417,"items":439},[440,444,447,450,453,456,459,462,465,468,471,474,477,480,483,486,489],[441],{"tag":442,"href":32,"children":443},"a",[31],[445],{"tag":442,"href":35,"children":446},[34],[448],{"tag":442,"href":38,"children":449},[37],[451],{"tag":442,"href":41,"children":452},[40],[454],{"tag":442,"href":44,"children":455},[43],[457],{"tag":442,"href":47,"children":458},[46],[460],{"tag":442,"href":50,"children":461},[49],[463],{"tag":442,"href":53,"children":464},[52],[466],{"tag":442,"href":56,"children":467},[55],[469],{"tag":442,"href":59,"children":470},[58],[472],{"tag":442,"href":62,"children":473},[61],[475],{"tag":442,"href":65,"children":476},[64],[478],{"tag":442,"href":68,"children":479},[67],[481],{"tag":442,"href":71,"children":482},[70],[484],{"tag":442,"href":74,"children":485},[73],[487],{"tag":442,"href":77,"children":488},[76],[490],{"tag":442,"href":80,"children":491},[79],[493,541,603,645],{"slug":494,"published":495,"minutes":6,"category":7,"tags":496,"keywords":502,"about":509,"sources":513,"cover":532,"og":533,"expertise":83,"locales":534,"lang":85,"title":535,"description":536,"coverAlt":537,"url":538,"pricing":539,"kind":540},"ollama","2026-09-29",[497,498,499,500,501],"Local inference","Open models","llama.cpp","GGUF","Model serving",[494,503,504,505,506,507,508],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[510],{"name":511,"url":512},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[514,517,520,523,526,529],{"title":515,"url":516},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":518,"url":519},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":521,"url":522},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":524,"url":525},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":527,"url":528},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":530,"url":531},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[85,86,87],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":542,"published":543,"minutes":544,"category":7,"tags":545,"keywords":551,"about":559,"sources":571,"cover":596,"og":597,"expertise":83,"locales":598,"lang":85,"title":599,"description":600,"coverAlt":601,"url":562,"pricing":602,"kind":546},"langfuse","2026-08-13",10,[546,547,548,549,550],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[542,552,553,554,555,556,557,558],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[560,563,565,568],{"name":561,"url":562},"Langfuse","https:\u002F\u002Flangfuse.com",{"name":548,"url":564},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":566,"url":567},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":569,"url":570},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[572,575,578,581,584,587,590,593],{"title":573,"url":574},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":576,"url":577},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":579,"url":580},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":582,"url":583},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":585,"url":586},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":588,"url":589},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":591,"url":592},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":594,"url":595},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[85,86,87],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":604,"published":605,"minutes":544,"category":7,"tags":606,"keywords":611,"about":618,"sources":621,"cover":637,"og":638,"expertise":83,"locales":639,"lang":85,"title":640,"description":641,"coverAlt":642,"url":643,"pricing":644,"kind":9},"openrouter","2026-07-23",[9,607,608,609,610],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[604,612,17,613,614,615,616,617],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[619],{"name":393,"url":620},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[622,625,626,629,632,635],{"title":623,"url":624},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":79,"url":80},{"title":627,"url":628},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":630,"url":631},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":633,"url":634},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":636,"url":620},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[85,86,87],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",{"slug":646,"published":647,"minutes":544,"category":7,"tags":648,"keywords":652,"about":659,"sources":667,"cover":691,"og":692,"expertise":83,"locales":693,"lang":85,"title":694,"description":695,"coverAlt":696,"url":662,"pricing":697,"kind":698},"braintrust","2026-07-07",[649,12,650,651,547],"Evaluations","LLM-as-a-judge","CI gates",[646,653,654,655,656,657,658],"braintrust pricing","braintrust eval","autoevals library","braintrust vs langfuse","llm evaluation platform","eval driven development",[660,663,666],{"name":661,"url":662},"Braintrust","https:\u002F\u002Fwww.braintrust.dev",{"name":664,"url":665},"Continuous integration","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FContinuous_integration",{"name":569,"url":570},[668,671,674,677,680,683,686,688],{"title":669,"url":670},"Braintrust pricing: plans, credits and usage rates","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing",{"title":672,"url":673},"Braintrust documentation: plans and limits","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fplans-and-limits",{"title":675,"url":676},"Braintrust documentation: evaluation quickstart","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fevaluation-quickstart",{"title":678,"url":679},"Braintrust documentation: get started","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs",{"title":681,"url":682},"GitHub: Braintrust organisation repositories","https:\u002F\u002Fgithub.com\u002Forgs\u002Fbraintrustdata\u002Frepositories",{"title":684,"url":685},"Arize Phoenix documentation: self-hosting","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix\u002Fself-hosting",{"title":687,"url":580},"Langfuse pricing: cloud plans and billable units",{"title":689,"url":690},"LangChain pricing: LangSmith plans","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fog.jpg",[85,86,87],"Braintrust review: eval-first observability with a hard meter","Braintrust turns production traces into datasets and gated experiments. What Starter and Pro really include, which parts are open source, and where Phoenix, Langfuse and LangSmith win.","A loop from instrumented application logs into a dataset, an experiment with scorers, and a comparison that gates the pull request before the change returns to the application.","Free · from $249 per month","Evaluation platform",1791383548682]