Tools/LLMOps & evals
LiteLLM: the OpenAI-compatible gateway most platform teams end up running
LiteLLM is the MIT-licensed OpenAI-compatible gateway most platform teams put in front of their providers. What it does well, what it costs and where it breaks.
- Type
- LLM gateway
- Pricing
- MIT · Enterprise from $20 per seat
Balázs Csorba··10 min read
- LLM gateway
- OpenAI-compatible
- Cost control
- Routing
- MCP

Key takeaways
- LiteLLM is an MIT-licensed Python SDK and proxy that puts 140+ providers behind one OpenAI-shaped endpoint, with virtual keys, per-team budgets and cross-provider fallbacks.
- The value is organisational rather than technical: changing model becomes a config edit, and every request produces a spend row.
- Overhead is measurable and published: the docs report 8ms p95 at 1,000 requests per second, and are unusually explicit that latency only means something next to the number of requests in flight.
- The free tier is the whole gateway. SSO beyond five users, audit logs, key rotation and the built-in moderation callbacks sit behind a licence quoted by annual request capacity, with no public list price.
- Release cadence is roughly one minor line a week against a four-line support window, so the upgrade is a recurring operational task rather than an event.
LiteLLM is the open-source LLM gateway that most platform teams end up running whether or not they meant to. It is two products in one package: a Python SDK that normalises the major providers onto the OpenAI request and response shape, and a proxy server that puts virtual keys, budgets, routing and spend tracking in front of them. On the documentation and the published numbers it is the most complete piece of plumbing in this category, and heavy enough that it deserves a deliberate decision rather than an accidental one.
It sits one hop in front of the model providers, not in front of the application data. Retrieval, prompt management and agent orchestration live elsewhere; an OpenAI client, a LangGraph workflow and a LiteLLM proxy coexist without complaint. What the gateway owns is the request itself: who is allowed to make it, which deployment serves it, what it costs and who is accountable for the bill.
What LiteLLM is
The name covers two runnable things that share a config file and a model price table. The SDK is a library you import. The proxy is a server that speaks the OpenAI wire format, so any client that already works against OpenAI can be redirected by changing its base URL and its key.
- MIT-licensed, with an Enterprise tier that adds identity and governance features rather than capacity.
- 140+ provider integrations and roughly 1,900 models, tracked in a public price and context-window file that ships with the package.
- OpenAI-compatible endpoints including chat completions, the Anthropic messages route, the responses API, embeddings, batches and realtime.
- Virtual keys with per-key, per-team and per-user budgets, rate limits and model allowlists, all backed by Postgres.
- A router with several deployment-selection strategies, per-deployment cooldowns, weighted failover, cross-model fallbacks and session affinity.
- An MCP gateway and an A2A agent gateway on the same endpoint and the same key policy, so agents get governed tool access rather than a separate one.
- Callbacks into Langfuse, MLflow, Helicone, LangSmith and OpenTelemetry, plus Prometheus metrics and an admin UI.
How it works
Every request through the proxy follows the same order of operations: authenticate the key, check the budget and rate limits, count tokens, pick a deployment, call the provider, translate the response, record the cost. Each step can fail on its own, and the router's behaviour on failure is where most of the operational learning happens.
Model names are the abstraction that makes this work. The config file lists deployments under a shared alias, several deployments may carry the same alias, and a request for that alias is answered by whichever deployment the routing strategy picks. Move the alias and every caller changes provider without a code change. The same table holds the fallbacks: a failed deployment is either re-picked from its peers or escalated to a named fallback model.
Measured overhead
The benchmark page reports 8ms p95 at 1,000 requests per second across four instances of 4 vCPU and 8 GB, and for the realtime endpoint 59ms median, 67ms p95 and 99ms p99 at 1,207 RPS. The same page publishes a head-to-head against Portkey at roughly 1,170 RPS on the same hardware, where LiteLLM reports p95 150ms and p99 240ms against Portkey's 230ms and 500ms.
The more useful part is the caveat the docs volunteer. The load generator sleeps 0.5 to 1 second between requests, so about 130 requests are in flight at any instant; a closed-loop client with no think time holds 1,000 and therefore reports roughly eight times the latency at the same throughput. A published overhead number is meaningless without the in-flight depth next to it, which is more disclosure than most vendor benchmarks carry.
Getting started
The smallest useful deployment is a config file with two deployments behind one alias, a master key and a Postgres connection. Caching, guardrails, callbacks and the admin UI all come later.
# config.yaml: two deployments behind one alias, plus a managed fallback
model_list:
- model_name: chat-fast
litellm_params:
model: openai/gpt-5.6-luna
api_key: os.environ/OPENAI_API_KEY
rpm: 900
- model_name: chat-fast
litellm_params:
model: anthropic/claude-sonnet-5
api_key: os.environ/ANTHROPIC_API_KEY
rpm: 60
- model_name: chat-cheap
litellm_params:
model: openai/gpt-5.6-luna
api_key: os.environ/OPENAI_API_KEY
router_settings:
routing_strategy: simple-shuffle
num_retries: 2
fallbacks:
- chat-fast: [chat-cheap]
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/DATABASE_URL
# docker run -p 4000:4000 -v $(pwd)/config.yaml:/app/config.yaml \
# -e DATABASE_URL -e LITELLM_MASTER_KEY \
# docker.litellm.ai/berriai/litellm:latest --config /app/config.yamlThe client change is two lines: point the OpenAI SDK at the proxy and pass a virtual key instead of the provider key. Every response carries an x-litellm-overhead-duration-ms header reporting what the gateway itself cost, which is the number to alert on when latency regresses and the upstream did not change.
Keys, budgets and spend
The key model is where LiteLLM earns its keep. A virtual key carries a model allowlist, a maximum budget with a reset window, TPM and RPM limits, a parallel request cap and an owner. Keys inherit from their owner, which is convenient and also the sharpest edge in the design: a key minted by an admin without a user id inherits nothing at all, and a key owned by a proxy admin can reach the admin management routes because management access follows the owner's role rather than the key's own permissions.
# mint a key for one team, capped, rate-limited and attributed
curl 'http://localhost:4000/key/generate' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H 'Content-Type: application/json' \
-d '{
"models": ["chat-fast"],
"max_budget": 40,
"budget_duration": "30d",
"rpm_limit": 600,
"tpm_limit": 400000,
"metadata": {"team": "platform", "env": "prod"}
}'
# what has this key spent, and against which cap
curl 'http://localhost:4000/key/info?key=sk-...' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY"Spend is computed from the same public price file the SDK ships, written on every completion, embedding and image call, and queryable per key, user and team. Those numbers are therefore estimates derived from a price table that can lag a provider's own billing: close enough to charge a team back, not close enough to reconcile an invoice. When a cost figure matters for a contract, take it from the provider invoice.
Pricing and licence boundaries
The gateway is free to self-host forever under MIT, with no per-seat and no per-token charge. The pricing page lists the open-source tier as 140+ provider integrations, virtual keys, users and teams, spend tracking, budgets and rate limits, LLM fallbacks, request and response logging, Prometheus metrics and guardrails. Procurement runs directly or through AWS Marketplace and resellers.
The licence buys control, not throughput. SSO is included up to five users and needs a licence beyond that; SCIM, audit logs, secret-manager write-back, IP allowlists, the multi-region control plane, virtual key rotation and the built-in moderation callbacks all sit behind it, and the docs name them individually so nobody discovers the boundary after rollout. Enterprise pricing is quoted against annual request capacity and deployment shape, never per token, and no list price is published, so any real cost comparison starts with a sales conversation.
Where it shingles
The weaknesses are operational rather than functional. The proxy needs Postgres and, once it runs on more than one instance, Redis; that is a stateful system with schema migrations and connection arithmetic, and the docs carry dedicated sizing pages precisely because it is where deployments fail. The Python SDK pulls a large dependency tree into a service for what is mostly request formatting. And the March 2026 supply-chain incident, in which trojanised LiteLLM releases reached PyPI, is a reminder that this is a package with very high install counts handling production keys.
| LiteLLM | Portkey | Bifrost | |
|---|---|---|---|
| Deployment | Self-hosted, air-gapped supported | Managed cloud and self-hosted | Self-hosted |
| Cost basis | Free OSS, licence by capacity | Per-month subscription | No published list price |
| Vendor p99 overhead | 0.66 ms | 2.29 ms | 4.54 ms |
| Best fit | Platform team owning access | Guardrails first | Minimal hop overhead |
Those overhead figures come from LiteLLM's own benchmark, run with every gateway pointed at the same deterministic mock upstream on identical hardware, which makes the comparison fair and the winner inevitable. For anything that is not an LLM gateway, the Kubernetes-native option deserves a look: Envoy AI Gateway exposes the OpenAI routing surface without adding a Python runtime to the request path, which is the structural difference that matters more than any of the numbers above.
Verdict
LiteLLM is the right default for a platform team that must give many teams access to several providers and needs three questions answered in production: who spent this, which deployment served it, and what happens when that provider degrades. It answers all three better than the alternatives, and it does so without taking custody of the provider keys.
- Adopt it when more than one team calls models through more than one provider. Budgets, spend attribution and cross-provider fallback are what justify the stateful deployment.
- Adopt it when procurement consolidation is the goal. Moving from one model to another becomes a config edit rather than a security review cycle.
- Take the SDK alone first if there is a single service. It normalises errors and responses for a fraction of the operational cost and leaves the door open to the proxy later.
- Skip it if one team and one provider is the whole story. A provider SDK and a spend spreadsheet are less machinery for the same result.
- Skip it if the last millisecond matters on the hot path. The gateway is a Python process in the request path, and the Rust rewrite is still labelled beta.
Sources
Frequently asked questions
Is LiteLLM just a wrapper around the OpenAI SDK?
No. The SDK normalises requests, responses and exceptions across the supported providers, but the proxy adds virtual keys, spend tracking, budgets, routing, cooldowns, fallbacks and caching on top. The usual path is to start with the SDK inside one service and move to the proxy once more than one team needs access.
How much latency does the gateway add?
The benchmark page reports 8ms p95 at 1,000 requests per second on four instances of 4 vCPU and 8 GB. That figure includes the client: the docs point out that with no think time a closed-loop client holds 1,000 requests in flight instead of about 130, which by Little's Law reports roughly eight times the latency at the same throughput.
Does the free version really enforce budgets?
Yes. Virtual keys, teams, spend tracking, budgets, rate limits and LLM fallbacks are all in the open-source tier. What the licence gates is identity and governance: SSO beyond five users, SCIM, audit logs, secret-manager write-back, virtual key rotation and the built-in moderation callbacks.
How is Enterprise priced?
There is no public list price. The pricing page says the licence is sized to annual gateway request capacity, deployment architecture and support needs, never per token, and every quote comes from sales. A 30-day trial key is issued without a credit card, and procurement is also available through AWS Marketplace and resellers.