Tools/LLMOps & evals

OpenRouter: one API key in front of every model you might call

OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.

Type
LLM gateway
Pricing
Pay per token, no subscription

··10 min read

  • LLM gateway
  • Model routing
  • Fallbacks
  • OpenAI-compatible
  • Pay per token
Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.

Key takeaways

  • OpenRouter is a hosted routing layer: one OpenAI-compatible endpoint in front of 500+ models from 80+ providers, billed at the provider list price with a 5.5% fee on credit purchases.
  • Default routing is price-weighted: among providers with no outage in the last 30 seconds, the router picks by the inverse square of the price and keeps the rest as fallbacks.
  • Failed and fallback attempts are not billed for model tokens, so failover costs latency rather than tokens.
  • The fee is taken when credits are bought, never per request, which makes the overhead a flat percentage you can budget instead of a line item you have to audit.
  • Zero data retention is available on every plan, but keeping prompts inside the EU or the US starts at the Business plan, where the platform fee is 8%.

OpenRouter is a hosted routing layer for model APIs: one OpenAI-compatible endpoint that forwards each request to whichever provider is serving the chosen model. It hosts no models of its own, and the pricing page states that inference is billed at the provider list price, with the platform fee charged when credits are bought. This review takes the view that it is the cheapest way to put a changing catalogue behind one key, and the most expensive dependency to remove once every service speaks its dialect.

It occupies the slot a self-hosted gateway would occupy, competing with LiteLLM, Portkey and the habit of calling each provider directly. The distinction is ownership rather than features. A gateway is a process you deploy, patch and scale; OpenRouter is a service in the request path of every call the product makes, so availability, data policy and pricing become somebody else’s problem, which is precisely what is being bought.

What it is

The company has routed traffic since 2023 and describes the product as a unified interface for hundreds of models. The documented surface is deliberately small: one endpoint, one key, a model slug and a few optional routing hints. Provider health, price comparison and failover all happen on the far side of that call.

  • Hosted only: there is no self-hosted edition, and the product is the endpoint at https://openrouter.ai/api/v1.
  • 500+ models from 80+ providers on the paid plans; the Free plan exposes 25+ free models across 4 providers.
  • OpenAI-compatible chat completions, plus an Anthropic messages route, a Responses API, embeddings and batches.
  • Model fallbacks through a models array: the next model is tried when the first errors, rate-limits or is refused by moderation.
  • Provider control per request — order, only, ignore, quantisation filters, and sorting by price, throughput or latency.
  • Bring-your-own-key routing, with the first $25,000 of list-price inference per month on a provider key charged at no OpenRouter fee.
  • Activity logs with export on every plan, per-generation records, and trace export to Datadog, Grafana Cloud, Langfuse, OpenTelemetry, Snowflake or a webhook.

How it works

The default strategy is price-based load balancing. The router first drops providers with a significant outage in the last 30 seconds, then picks among the stable ones weighted by the inverse square of the price, keeping the remainder as fallbacks: between a $1 endpoint and a $3 endpoint the cheap one is nine times more likely to be tried first. Any explicit sort or order disables that balancing, which is the switch that turns a cost-optimising router into a predictable one.

OpenRouter request pathAn application sends a chat completion to one OpenAI-compatible endpoint. The router sorts candidate providers by price, the fallback list names the next model, a provider answers, and the response reports the model that actually served. Workspaces, guardrails and activity logs sit underneath the request path.OpenRouter request pathopenai-compatibleClientOpenAI SDKRoutersort, orderFallbackmodels[] listProvider80+ vendorsResponsemodel fieldCONTROL PLANEWorkspacesbudgets, keysGuardrailsspend, policyActivity logsexport, tracesThe fee is taken at the credit purchase, not on the request
One endpoint replaces a rack of provider integrations. Routing, failover and billing all live on the far side of the call, which is exactly what the platform fee buys.

The cost of the hop

Every call now crosses two networks instead of one. The pricing FAQ states plainly that routing improves reliability while latency varies by model, provider and region, and its advice for consistent latency is to pin a model and a provider. The extra hop is negligible next to a generation that runs for seconds and is not negligible next to a sub-second classification call: pinning returns the behaviour of a direct API plus one round trip.

The indirection stays visible where it matters. The response names the endpoint that actually served the request, and an opt-in header surfaces the routing decision on every response, so an unexpected provider shows up in a log line rather than in a customer complaint.

  • sort: price, throughput or latency, and it switches load balancing off.
  • order: an explicit provider sequence, with allow_fallbacks false when nothing outside it may be tried.
  • require_parameters: only providers that support every parameter in the request, which stops structured outputs and tool calls from being quietly degraded.
  • zdr and data_collection: route only to zero-data-retention endpoints, or only to providers that do not store inputs.

Getting started

One key and one POST. Free accounts get a working endpoint immediately: free models are capped at 20 requests per minute and 50 per day, and the daily allowance rises to 1,000 once the account has bought at least $10 of credits.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "models": ["~openai/gpt-sol-latest", "~anthropic/claude-sonnet-latest"],
    "messages": [
      { "role": "user", "content": "Summarise this stack trace in one sentence." }
    ],
    "provider": {
      "sort": "price",
      "require_parameters": true
    }
  }'

The answer comes back in the OpenAI shape, and the model field names the endpoint that actually served it. That is also the model the request is billed at, whether or not it was first in the list. An existing OpenAI client needs only a different base URL; the documentation shows the same swap in Python and TypeScript.

Pricing

There is no subscription, no minimum and no lock-in. Credits are bought upfront by card, AliPay or USDC, and each request deducts the provider list price from the balance; the platform fee, 5.5% on Standard and 8% on Business, is charged on the purchase and never on the request.

  • Free: 25+ free models, 4 providers, activity logs with export, 50 requests per day on free models. No auto-routing, no budgets, no prompt caching.
  • Standard: the full 500+ models and 80+ providers at a 5.5% fee on credit purchases, with auto-routing, budgets, prompt caching, management API keys and five workspaces.
  • Business: the same catalogue at 8%, EU or US in-region routing through eu.openrouter.ai and us.openrouter.ai, 1,000 workspaces and workload identity federation.
  • Enterprise: contracted fees, invoicing, SSO and SCIM, contractual SLAs, and $200,000 of BYOK list-price inference per month before a 5% charge applies.

Control plane

For teams, the parts that matter are not in the request body. Workspaces, budgets and guardrails decide who may call what with which key, and they live in the platform rather than in your code, which is the point: policy changes stop being deployments.

  • Workspaces separate environments: five on Free and Standard, 1,000 on Business, each with its own keys and limits, plus a credit cap per key.
  • Guardrails attach to keys or members and cover spend, model access, prompt-injection patterns and sensitive-information rules.
  • Management API keys create, list and delete keys programmatically, so key lifecycle belongs in CI rather than in a dashboard habit.
  • Broadcast ships traces to Datadog, Grafana Cloud, Langfuse, OpenTelemetry, Sentry, Snowflake or any HTTP endpoint, which is what makes spend joinable with the rest of the stack.

Zero data retention runs on every plan, account-wide or per request with zdr. Regional residency is the feature that is not free: keeping prompts and completions inside the EU or the US starts at Business, which makes the 8% plan the baseline rather than the upgrade for regulated workloads.

Where it shingles

The weaknesses are structural. Every request depends on a third party’s availability, and the routing decision stays invisible unless the metadata header is enabled. Parameter support differs between providers of the same model, so a call that works through one endpoint can be quietly degraded through another unless require_parameters is set. Provider price changes flow straight through — the FAQ states that you will be charged at the new rate — and the free tier is rate-limited rather than generous. There is also ownership: Bloomberg and the Wall Street Journal reported in August 2026 that Stripe had agreed to buy the company, a useful reminder that the layer inside the request path has an owner.

OpenRouterLiteLLMFirst-party APIs
DeploymentHosted, no self-hosted editionSelf-hosted proxy you operateYour code per provider
Fee5.5% on credits (Standard)Free under MIT, paid tiers by quoteNone
Catalogue500+ models, 80+ providers140+ providers, ~1,900 modelsOne vendor per integration
FallbackPer request, cross-provider and cross-modelConfigured in the routerWritten by hand
Best fitMany models, one bill, no operationsPlatform team owning keys and budgetsOne model on a hot path

The comparison that matters is custody rather than features. OpenRouter removes the operations and adds a percentage plus a dependency; LiteLLM keeps the keys and the workload inside your infrastructure; first-party APIs give the shortest path and the least cover when a model is unavailable. For a product that changes models as often as others change a pricing page, that cover is worth more than the fee.

Verdict

OpenRouter answers how to give many agents, services and teams access to a fast-changing model list without operating a gateway. It does not answer how to own the request path.

  1. Adopt it when the model list changes faster than the integration: one endpoint and a models array turn a vendor evaluation into a configuration value.
  2. Adopt it for agent and coding-agent workloads, where the bill spans many models and one activity log with export beats four provider dashboards.
  3. Use BYOK when contracts or data policy require a direct relationship with the provider; routing and analytics are kept, and the first $25,000 of list-price inference per month is fee-free.
  4. Skip it on latency-critical paths that cannot pin a provider: availability routing is the product, and a request that may change endpoint is a tail latency nobody controls.
  5. Skip it when one provider and one model is the whole product: a first-party SDK, direct billing and no intermediary percentage is less machinery for the same call.
“Inference is billed at the provider’s list price on every plan.” the pricing page The fee is small and legible; the dependency is the part that has to be priced in, and it is not listed anywhere.

Sources

  1. OpenRouter documentation: quickstart
  2. OpenRouter pricing
  3. OpenRouter documentation: model fallbacks
  4. OpenRouter documentation: provider routing
  5. OpenRouter documentation index
  6. Wikipedia: OpenRouter

Frequently asked questions

Does OpenRouter mark up model prices?

The pricing page says no: inference is billed at the provider list price on every plan, and the platform fee of 5.5% on Standard and 8% on Business is charged when credits are purchased. Routing through openrouter/auto adds no fee either, so you pay the rate of whichever model serves the request.

How much latency does the gateway add?

It adds one network hop before the provider, and the documentation states plainly that routing improves reliability while latency varies by model, provider and region. Its advice for consistent latency is to pin a model and a provider, which switches load balancing off and makes the path behave like a direct API call.

What happens when a fallback triggers?

Any error can trigger the next model in the models array, including context-length validation, moderation flags, rate limits and downtime. The attempt that errors is not charged for its model tokens, so the bill covers the run that answers, but the caller waits for both.

Can prompts be kept inside one region?

Yes, on Business and Enterprise: requests to eu.openrouter.ai or us.openrouter.ai keep prompts and completions inside that region, with no cross-region fallback. Zero data retention is available on every plan, account-wide or per request.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.