Tools/LLMOps & evals

Portkey: a production LLM gateway, reviewed for routing, guardrails and cost

Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.

Type
LLM gateway
Pricing
Free · from $49 per month

··11 min read

  • LLM gateway
  • Guardrails
  • Routing
  • Observability
  • Cost control
A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.

Key takeaways

  • Portkey's routing config is the strongest part of the product: a closed schema, four nestable strategy modes, weights that normalise to 100 and a circuit breaker whose cooldown cannot be set below 30 seconds.
  • The hosted gateway is a network hop. Portkey's own benchmark repository measures it at plus 93 ms on average against a direct Bedrock call, and calls 50 to 150 ms typical for the extra two hops.
  • Guardrails only ever read the last message and never read image inputs, and a denied synchronous check returns 446, a status code most SDKs will retry as if it were a transport error.
  • Pricing is per recorded log rather than per request, and the pricing page and the caching documentation disagree about whether the 49 dollar Production plan includes semantic caching.
  • Since the May 2026 acquisition it ships as the gateway inside Prisma AIRS, which changes who signs the contract without changing the API.

Portkey is an LLM gateway: a proxy that sits between an application and the model providers it calls, and turns provider keys, retries, fallbacks, caching, guardrails and cost attribution into configuration objects rather than application code. It is one of the more complete gateways on the market, and since Palo Alto Networks closed its acquisition of Portkey in May 2026 it also ships as the gateway inside Prisma AIRS. The position taken here: the routing layer is the strongest in its class and worth the operational dependency, while the hosted deployment and the guardrail semantics are where the caveats sit.

What it replaces is the hand-rolled retry wrapper every team writes in the first month of an LLM product, which is usually a try block with a sleep, one hard-coded provider key and no record of what any of it cost. Portkey moves that behind a single OpenAI-compatible base URL, so the OpenAI SDK, the Anthropic SDK, LangChain or a raw fetch call all talk to the same endpoint. It competes directly with LiteLLM, free and self-hosted; with OpenRouter, a paid marketplace rather than a control plane; and with Cloudflare AI Gateway, which is close to free because it rides infrastructure most teams already pay for.

What it is

Portkey is really two products sharing one repository. The gateway itself is an MIT-licensed Node proxy that runs from a single command and listens on port 8787 with a local console attached; the control plane around it is a paid service that stores credentials, hosts the configuration UI and keeps the logs. That split explains most of what follows. What the proxy does is well documented and free, and what the control plane does is where the pricing, the guardrail catalogue and the compliance story live.

  • Runs as one command, npx @portkey-ai/gateway, serving on localhost:8787 with a console at /public/.
  • MIT licensed with roughly 13,100 GitHub stars and 1,300 forks; the published npm package sits at version 1.15.2 and a 2.0 pre-release branch is in progress.
  • One OpenAI-compatible endpoint, plus Anthropic's /v1/messages and the Open Responses format, so the model string is the only thing that changes when the provider does.
  • Model strings carry the provider: @openai-prod/gpt-4o resolves to a stored integration, along with its budget and its rate limit.
  • Strategy objects nest. A fallback can contain a load balancer that contains another fallback, each with its own weights, status-code triggers and circuit breaker.
  • Guardrails evaluate inputs and outputs, running either asynchronously with no added latency or synchronously with documented deny codes of 246 and 446.
  • Owned by Palo Alto Networks since May 2026 and sold as the Prisma AIRS AI Gateway, generally available since 16 July 2026.

How it works

A request arrives carrying a Portkey API key and, usually, a config. The config names a strategy and a list of targets. The gateway resolves each target to a stored provider integration, applies caches and guardrails as the config dictates, forwards to the chosen provider, and writes a log line with latency, token counts and cost. Where this differs from a hand-rolled proxy is not the forwarding. It is that the decision is data: the same JSON runs unchanged in the hosted product, in the open-source proxy and in a self-hosted data plane.

A request through the Portkey gatewayAn application sends an OpenAI-compatible request to the Portkey gateway. The gateway applies the attached config, then fans the request out to one of three providers using retry, fallback and weighted routing. A guardrail verdict is checked synchronously and returns 246 or 446. Every call writes a log line with latency, tokens and cost.your appOpenAI SDKPortkey gatewayconfig, retry, cacheguardrail, logguardrail246, or 446logscost, latencyOpenAIAnthropicBedrockretry, fallbacksyncevery call
One request through the gateway: the config decides the target, a synchronous guardrail can deny with 246 or 446, and every call is logged with its latency, tokens and cost.

The interesting branch is the guardrail. Run asynchronously, which is the default, the check runs alongside the model call, the result is only logged, and the provider's own status code comes back unchanged. Run synchronously, the check blocks: a pass returns 200, a failure returns 246 when deny is off and 446 when it is on. Both codes sit outside the range any client library was written against. A library that treats any status other than 200 as success will silently accept a 246 that was supposed to be flagged, and one that treats anything outside the 2xx range as an exception will throw on a 446 and retry it as though the network had failed. That is the sharpest edge in the product.

The config object is the other half. Its schema is closed, with a fixed enum of four strategy modes, single, loadbalance, fallback and conditional, and a fixed set of keys, so a typo is rejected rather than silently ignored. Targets are themselves configs, which is what makes the strategies compose.

{
  "strategy": { "mode": "fallback", "on_status_codes": [429, 500, 503] },
  "retry": { "attempts": 3, "use_retry_after_headers": true },
  "cb_config": { "failure_threshold": 5, "cooldown_interval": 60000 },
  "targets": [
    { "provider": "@openai-prod",
      "override_params": { "model": "gpt-4o" } },
    { "strategy": { "mode": "loadbalance" },
      "targets": [
        { "provider": "@anthropic-prod", "weight": 0.8,
          "override_params": { "model": "claude-sonnet-4-5-20250929" } },
        { "provider": "@bedrock-prod", "weight": 0.2,
          "override_params": { "model": "anthropic.claude-3-5-sonnet-20241022-v2:0" } }
      ] }
  ]
}

Three details earn their own attention. Weights are normalised to 100 and a weight of zero keeps a target in the config while sending it no traffic, which is how a canary is paused rather than deleted. The circuit breaker's cooldown_interval has a floor of 30,000 milliseconds, so a fast-flap protection loop cannot be configured into a tight retry storm. And sticky routing hashes the fields you name with a one-hour default TTL, but its two-tier cache is in-memory plus Redis, so without Redis it works on a single instance only.

Getting started

Install the SDK, add a provider in the Model Catalog, and change one string. The Portkey SDK is a superset of the OpenAI client, so an existing integration usually needs nothing but its base URL and key swapped.

from portkey_ai import Portkey

client = Portkey(api_key="PORTKEY_API_KEY")

answer = client.chat.completions.create(
    model="@openai-prod/gpt-4o",              # @provider-slug/model-name
    messages=[{"role": "user", "content": "Summarise this ticket in one line."}],
)
print(answer.choices[0].message.content)

# Same code, different model: only the string changes
answer = client.chat.completions.create(
    model="@anthropic-prod/claude-sonnet-4-5-20250929",
    max_tokens=512,
    messages=[{"role": "user", "content": "Summarise this ticket in one line."}],
)
print(answer.choices[0].message.content)

Two things matter before this goes near production. The provider slug is resolved server-side from the Model Catalog rather than from a literal in the request, so a wrong model string fails at request time and not at deploy time. And the config that carries retries, caching and guardrails is not in the snippet: it is attached either as a config ID on the client or as a JSON blob in the x-portkey-config header, which is exactly what lets routing policy change without a deploy.

Guardrails

Guardrails are checks attached to a request and evaluated on the input, the output or both. They are the part of Portkey most likely to be misconfigured, partly because the feature list reads wider than the behaviour. The documentation is unusually explicit about the limits, which at least makes the limits easy to find.

  • Only the last message in the request is evaluated, and only its text portions. Image inputs, base64 or URL, are not checked at all.
  • Guardrails do not run on the Assistants, Audio, Images, Files, Batch, Fine-tuning, Moderations or Models endpoints. They do run on chat completions, completions, embeddings with input only, messages, responses and prompt completions.
  • Output guardrails on a streamed response are informational: the verdict arrives as a trailing chunk after the done marker and triggers no fallback and no retry.
  • To see hook results in a stream at all, strict OpenAI compliance has to be switched off with the x-portkey-strict-open-ai-compliance header, because the default strips them.
  • Tiers follow the plan: basic checks on Developer, basic plus partner and pro checks on Production, everything including custom on Enterprise. Most checks are deterministic, regex, JSON schema, word and character counts, with LLM-based checks such as prompt-injection scanning on top.
  • Partner guardrails exist, including Aporia, Pillar Security, SydeLabs, Zscaler AI Guard and Akto, each called over HTTP with its own timeout: 10,000 ms for Zscaler, 5,000 ms for Akto.

The Anthropic path carries one more wrinkle. On /v1/messages the hook results arrive as a dedicated event the Anthropic SDK does not parse, so reading them means dropping down to cURL. That is the kind of small inconsistency that decides whether a feature gets adopted or quietly ignored, and it is worth checking against your own SDK before committing to a guardrail-based control.

Pricing and latency

Latency is where the marketing and the measurement diverge. Three numbers exist for the same product and they are not describing the same thing, which is why the table below names the source of each rather than picking the flattering one.

FigureSourceWhat it measures
Under 1 msGateway repository READMEThe self-hosted proxy's own processing time, not the hosted round trip
Sub-10 ms at 99.9999% uptimeVendor blog, October 2025A hosting claim covering 10 billion requests a month, with no published method
Plus 93 ms average, plus 25 ms medianPortkey's own benchmark repositoryThe cloud gateway against a direct Bedrock call, two workers, three requests per iteration
50 to 150 ms typicalThe same repository, overhead sectionWhat the vendor itself calls the round-trip cost of the two extra network hops

The honest reading is that the two figures describe two different products. Under 1 ms is the open-source proxy's processing time, which is genuinely good and is the main argument for self-hosting it. Sub-10 ms is a hosting claim with no method attached. The plus 93 ms figure is the only one published with a runnable harness, and the same repository describes 50 to 150 ms as typical. Against a 400 ms time to first token that is invisible; against a 90 ms autocomplete or voice path it is the entire budget. Measure it on your own traffic before adopting it, not from the pricing page.

Pricing follows, and it is charged on recorded logs rather than on requests. That is unusual and mostly harmless: the free plan keeps serving traffic after the log cap and simply stops recording. The paid tiers are where it starts to matter, because overage is priced per 100,000 requests on top of a log allowance.

PlanPriceRecorded logsRetention and what is included
DeveloperFree10,000 a month3 days for logs, 30 for metrics; 3 prompt templates; deterministic guardrails; community support; requests keep flowing past the cap
Production49 dollars a month100,000, then 9 dollars per extra 100,00030 days for logs, 90 for metrics; role-based access, service account keys, LLM and partner guardrails, semantic caching, production support
EnterpriseQuoted10 million or moreCustom retention; private cloud and VPC hosting, SSO, data export, SOC 2 Type 2, GDPR and HIPAA, data isolation

Enterprise is where the product becomes a different thing: a data plane inside your own VPC, Helm charts on Kubernetes 1.20 or later, one to two cores and two to four gigabytes of memory per instance, logs in S3-compatible object storage or MongoDB, and a control plane the gateway syncs from once a minute while holding a seven-day local cache of configs and keys. The documentation recommends a volatile-lru eviction policy so that live configuration survives memory pressure. This is a deployment to operate, not a container to forget about.

Where it does not fit

The honest weaknesses first, because they rule the tool out for whole classes of team. A hosted gateway is a single point of failure and a single tenancy boundary for every request your application makes, and the status page does show control plane incidents, including two in August 2026 lasting 40 minutes and two hours. Semantic caching is gated behind an Enterprise conversation, so the cost-saving feature teams cite in their business case is not available to them. The guardrail model only ever sees the last message, which means it is not a prompt-injection defence for a multimodal conversation. And the config lives in a SaaS dashboard, which means the routing policy of a system is no longer reviewable in the repository that owns the system.

ToolLicence and costWhere it runsWhat it does not do
PortkeyMIT gateway, paid control plane from 49 dollarsVendor cloud, or your own VPC on EnterpriseNever reads image inputs; semantic cache is Enterprise-only on the hosted plan
LiteLLMMIT, self-host free, Enterprise priced by request capacityYour infrastructure onlyNo hosted tier, so no managed dashboards, shared guardrails or credential vault
OpenRouterPay per token, 5.5 percent platform fee on credit purchasesTheir cloudNo self-hosted gateway; BYOK runs above 25,000 dollars of list-price inference a month
Cloudflare AI GatewayCore features free on every plan, guardrails billed as Workers AI inferenceCloudflare edge, needs the Workers paid plan at volumeGuardrail checks are Llama Guard 3 8B on Workers AI, with no bring-your-own

The choice is therefore less about features than about where you want the policy to live. If the routing rules belong in version control next to the code, LiteLLM wins on every axis including cost. If you want a security team to be able to change them, if you want credentials that never touch an application environment, or if you are running inside a regulated environment that already buys from Palo Alto Networks, Portkey's control plane is the argument. Cloudflare AI Gateway is the third option worth pricing: core features are free on every plan and the cost is inference rather than platform, which inverts the calculation whenever traffic is small.

Verdict

Portkey is the most complete managed LLM gateway available, and the config object alone is a better design than most hand-built equivalents: nestable strategies, a closed schema, a hard floor on the circuit breaker cooldown. Take it when routing policy is a governance problem rather than a coding problem, and accept both the network hop and the fact that your routing config now lives somewhere other than your repository. The opinionated part is that this is the right default for enterprises and the wrong default for a small team, which would be better served by the free tier of the open-source proxy or by LiteLLM and a weekend of wiring.

  1. Take Portkey when several teams share model credentials and someone has to own budgets, allow-lists and rate limits centrally. That is the product's real job.
  2. Take it when a security team has to be able to change guardrails and inspect full request logs without a deploy.
  3. Do not take it for latency-sensitive completion or voice paths until you have measured the hop on your own traffic, because the published overhead is tens of milliseconds and the marketing figure is not.
  4. Do not rely on guardrails as a prompt-injection control for multimodal traffic. They never read the image, and they only read the last message.
  5. Skip it if your routing policy has to be reviewed in code review. Keep that config in the repository and run LiteLLM or the open-source gateway instead.
  6. Reconsider at the point where the control plane becomes the bottleneck: at that volume, a gateway on Cloudflare or an in-house proxy in your own VPC is cheaper and simpler than the enterprise tier.

Sources

  1. Portkey docs: AI Gateway
  2. Portkey docs: Getting started with the AI Gateway
  3. Portkey docs: Gateway config object
  4. Portkey docs: Guardrails
  5. Portkey docs: Guardrail endpoints and capabilities
  6. Portkey docs: Cache, simple and semantic
  7. Portkey docs: Load balancing
  8. Portkey docs: Enterprise hybrid deployment architecture
  9. Portkey pricing
  10. Portkey gateway on GitHub, MIT licensed
  11. Portkey's own benchmark: gateway versus direct Bedrock
  12. Portkey status page
  13. Palo Alto Networks completes acquisition of Portkey, May 2026
  14. Palo Alto Networks: Prisma AIRS AI Gateway
  15. Cloudflare AI Gateway pricing
  16. LiteLLM pricing
  17. OpenRouter pricing

Frequently asked questions

What is Portkey and what does an LLM gateway actually do?

An LLM gateway is a proxy that sits between your application and the model providers it calls. Portkey's turns provider keys, retries, fallbacks, weighted routing, caching, guardrails and cost attribution into JSON configuration objects attached to a request, so an OpenAI-compatible client keeps working while the routing policy changes in a dashboard instead of a deploy.

Does Portkey add latency to model calls?

On the hosted product, yes. Portkey's own benchmark repository measures an average of plus 93 ms and a median of plus 25 ms routing through the cloud gateway against a direct Bedrock call, and its README describes 50 to 150 ms as typical for the extra two hops. Self-hosted it is a different number: the gateway repository claims under 1 ms of processing.

Is Portkey open source?

The gateway is MIT licensed and runs from a single npm command, and the published package currently sits at version 1.15.2 with a 2.0 pre-release branch in progress. The control plane, the dashboard, the Model Catalog and the enterprise deployment options are commercial. Running the proxy yourself does not give you the managed product.

Portkey or LiteLLM?

LiteLLM is MIT licensed with no managed tier, so it wins on cost and on data staying inside your own network, and you build the dashboards yourself. Portkey wins when you want routing policy, guardrails and stored credentials to live in a product someone else operates, and you accept 49 dollars a month for 100,000 recorded logs plus 9 dollars per additional 100,000, along with a network hop.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.