[{"data":1,"prerenderedAt":674},["ShallowReactive",2],{"tool-openrouter-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":26,"cover":44,"og":45,"expertise":46,"locales":47,"lang":48,"title":51,"description":52,"coverAlt":53,"url":54,"pricing":55,"kind":9,"metaTitle":56,"takeaways":57,"faq":63,"toc":76,"blocks":101,"others":430},"openrouter","2026-07-23",10,"llmops",[9,10,11,12,13],"LLM gateway","Model routing","Fallbacks","OpenAI-compatible","Pay per token",[4,15,16,17,18,19,20,21],"openrouter vs litellm","llm gateway comparison","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[23],{"name":24,"url":25},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[27,30,33,36,39,42],{"title":28,"url":29},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":31,"url":32},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing",{"title":34,"url":35},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":37,"url":38},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":40,"url":41},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":43,"url":25},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg","ai-engineer",[48,49,50],"en","de","hu","OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription","OpenRouter review: one key, every model · Balázs Csorba",[58,59,60,61,62],"OpenRouter is a hosted routing layer: one OpenAI-compatible endpoint in front of 500+ models from 80+ providers, billed at the provider list price with a 5.5% fee on credit purchases.","Default routing is price-weighted: among providers with no outage in the last 30 seconds, the router picks by the inverse square of the price and keeps the rest as fallbacks.","Failed and fallback attempts are not billed for model tokens, so failover costs latency rather than tokens.","The fee is taken when credits are bought, never per request, which makes the overhead a flat percentage you can budget instead of a line item you have to audit.","Zero data retention is available on every plan, but keeping prompts inside the EU or the US starts at the Business plan, where the platform fee is 8%.",[64,67,70,73],{"q":65,"a":66},"Does OpenRouter mark up model prices?","The pricing page says no: inference is billed at the provider list price on every plan, and the platform fee of 5.5% on Standard and 8% on Business is charged when credits are purchased. Routing through openrouter\u002Fauto adds no fee either, so you pay the rate of whichever model serves the request.",{"q":68,"a":69},"How much latency does the gateway add?","It adds one network hop before the provider, and the documentation states plainly that routing improves reliability while latency varies by model, provider and region. Its advice for consistent latency is to pin a model and a provider, which switches load balancing off and makes the path behave like a direct API call.",{"q":71,"a":72},"What happens when a fallback triggers?","Any error can trigger the next model in the models array, including context-length validation, moderation flags, rate limits and downtime. The attempt that errors is not charged for its model tokens, so the bill covers the run that answers, but the caller waits for both.",{"q":74,"a":75},"Can prompts be kept inside one region?","Yes, on Business and Enterprise: requests to eu.openrouter.ai or us.openrouter.ai keep prompts and completions inside that region, with no cross-region fallback. Zero data retention is available on every plan, account-wide or per request.",[77,80,83,86,89,92,95,98],{"id":78,"title":79},"what-it-is","What it is",{"id":81,"title":82},"how-it-works","How it works",{"id":84,"title":85},"getting-started","Getting started",{"id":87,"title":88},"pricing","Pricing",{"id":90,"title":91},"control-plane","Control plane",{"id":93,"title":94},"where-it-shingles","Where it shingles",{"id":96,"title":97},"verdict","Verdict",{"id":99,"title":100},"sources","Sources",[102,106,109,112,115,153,154,164,173,177,180,183,221,222,225,227,234,249,250,253,270,276,277,280,290,296,297,303,359,362,379,380,383,399,409,410],{"type":103,"content":104},"paragraph",[105],"OpenRouter is a hosted routing layer for model APIs: one OpenAI-compatible endpoint that forwards each request to whichever provider is serving the chosen model. It hosts no models of its own, and the pricing page states that inference is billed at the provider list price, with the platform fee charged when credits are bought. This review takes the view that it is the cheapest way to put a changing catalogue behind one key, and the most expensive dependency to remove once every service speaks its dialect.",{"type":103,"content":107},[108],"It occupies the slot a self-hosted gateway would occupy, competing with LiteLLM, Portkey and the habit of calling each provider directly. The distinction is ownership rather than features. A gateway is a process you deploy, patch and scale; OpenRouter is a service in the request path of every call the product makes, so availability, data policy and pricing become somebody else’s problem, which is precisely what is being bought.",{"type":110,"level":111,"id":78,"text":79},"heading",2,{"type":103,"content":113},[114],"The company has routed traffic since 2023 and describes the product as a unified interface for hundreds of models. The documented surface is deliberately small: one endpoint, one key, a model slug and a few optional routing hints. Provider health, price comparison and failover all happen on the far side of that call.",{"type":116,"ordered":117,"items":118},"list",false,[119,126,128,130,136,149,151],[120,121,125],"Hosted only: there is no self-hosted edition, and the product is the endpoint at ",{"tag":122,"children":123},"code",[124],"https:\u002F\u002Fopenrouter.ai\u002Fapi\u002Fv1",".",[127],"500+ models from 80+ providers on the paid plans; the Free plan exposes 25+ free models across 4 providers.",[129],"OpenAI-compatible chat completions, plus an Anthropic messages route, a Responses API, embeddings and batches.",[131,132,135],"Model fallbacks through a ",{"tag":122,"children":133},[134],"models"," array: the next model is tried when the first errors, rate-limits or is refused by moderation.",[137,138,141,142,141,145,148],"Provider control per request — ",{"tag":122,"children":139},[140],"order",", ",{"tag":122,"children":143},[144],"only",{"tag":122,"children":146},[147],"ignore",", quantisation filters, and sorting by price, throughput or latency.",[150],"Bring-your-own-key routing, with the first $25,000 of list-price inference per month on a provider key charged at no OpenRouter fee.",[152],"Activity logs with export on every plan, per-generation records, and trace export to Datadog, Grafana Cloud, Langfuse, OpenTelemetry, Snowflake or a webhook.",{"type":110,"level":111,"id":81,"text":82},{"type":103,"content":155},[156,157,160,161,163],"The default strategy is price-based load balancing. The router first drops providers with a significant outage in the last 30 seconds, then picks among the stable ones weighted by the inverse square of the price, keeping the remainder as fallbacks: between a $1 endpoint and a $3 endpoint the cheap one is nine times more likely to be tried first. Any explicit ",{"tag":122,"children":158},[159],"sort"," or ",{"tag":122,"children":162},[140]," disables that balancing, which is the switch that turns a cost-optimising router into a predictable one.",{"type":165,"attrs":166,"inner":170,"caption":171},"diagram",{"viewBox":167,"role":168,"aria-labelledby":169},"0 0 720 320","img","or-t or-d","\u003Ctitle id=\"or-t\">OpenRouter request path\u003C\u002Ftitle>\u003Cdesc id=\"or-d\">An application sends a chat completion to one OpenAI-compatible endpoint. The router sorts candidate providers by price, the fallback list names the next model, a provider answers, and the response reports the model that actually served. Workspaces, guardrails and activity logs sit underneath the request path.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">OpenRouter request path\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">openai-compatible\u003C\u002Ftext>\u003Crect x=\"20\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"78\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Client\u003C\u002Ftext>\u003Ctext x=\"78\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">OpenAI SDK\u003C\u002Ftext>\u003Cpath d=\"M136 102 H148\" class=\"d-line\" \u002F>\u003Cpath d=\"M156 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"156\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"214\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Router\u003C\u002Ftext>\u003Ctext x=\"214\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">sort, order\u003C\u002Ftext>\u003Cpath d=\"M272 102 H284\" class=\"d-line\" \u002F>\u003Cpath d=\"M292 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"292\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"350\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Fallback\u003C\u002Ftext>\u003Ctext x=\"350\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">models[] list\u003C\u002Ftext>\u003Cpath d=\"M408 102 H420\" class=\"d-line\" \u002F>\u003Cpath d=\"M428 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"428\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"486\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Provider\u003C\u002Ftext>\u003Ctext x=\"486\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">80+ vendors\u003C\u002Ftext>\u003Cpath d=\"M544 102 H556\" class=\"d-line\" \u002F>\u003Cpath d=\"M564 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"564\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"622\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Response\u003C\u002Ftext>\u003Ctext x=\"622\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">model field\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"172\" class=\"d-label\">CONTROL PLANE\u003C\u002Ftext>\u003Cpath d=\"M214 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M214 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M368 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M368 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M546 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M546 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"134\" y=\"216\" width=\"160\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"214\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Workspaces\u003C\u002Ftext>\u003Ctext x=\"214\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">budgets, keys\u003C\u002Ftext>\u003Crect x=\"288\" y=\"216\" width=\"160\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"368\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Guardrails\u003C\u002Ftext>\u003Ctext x=\"368\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">spend, policy\u003C\u002Ftext>\u003Crect x=\"446\" y=\"216\" width=\"200\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"546\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Activity logs\u003C\u002Ftext>\u003Ctext x=\"546\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">export, traces\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"306\" text-anchor=\"middle\" class=\"d-label\">The fee is taken at the credit purchase, not on the request\u003C\u002Ftext>",[172],"One endpoint replaces a rack of provider integrations. Routing, failover and billing all live on the far side of the call, which is exactly what the platform fee buys.",{"type":110,"level":174,"id":175,"text":176},3,"latency-cost","The cost of the hop",{"type":103,"content":178},[179],"Every call now crosses two networks instead of one. The pricing FAQ states plainly that routing improves reliability while latency varies by model, provider and region, and its advice for consistent latency is to pin a model and a provider. The extra hop is negligible next to a generation that runs for seconds and is not negligible next to a sub-second classification call: pinning returns the behaviour of a direct API plus one round trip.",{"type":103,"content":181},[182],"The indirection stays visible where it matters. The response names the endpoint that actually served the request, and an opt-in header surfaces the routing decision on every response, so an unexpected provider shows up in a log line rather than in a customer complaint.",{"type":116,"ordered":117,"items":184},[185,199,207,212],[186,188,189,141,192,160,195,198],{"tag":122,"children":187},[159],": ",{"tag":122,"children":190},[191],"price",{"tag":122,"children":193},[194],"throughput",{"tag":122,"children":196},[197],"latency",", and it switches load balancing off.",[200,202,203,206],{"tag":122,"children":201},[140],": an explicit provider sequence, with ",{"tag":122,"children":204},[205],"allow_fallbacks"," false when nothing outside it may be tried.",[208,211],{"tag":122,"children":209},[210],"require_parameters",": only providers that support every parameter in the request, which stops structured outputs and tool calls from being quietly degraded.",[213,216,217,220],{"tag":122,"children":214},[215],"zdr"," and ",{"tag":122,"children":218},[219],"data_collection",": route only to zero-data-retention endpoints, or only to providers that do not store inputs.",{"type":110,"level":111,"id":84,"text":85},{"type":103,"content":223},[224],"One key and one POST. Free accounts get a working endpoint immediately: free models are capped at 20 requests per minute and 50 per day, and the daily allowance rises to 1,000 once the account has bought at least $10 of credits.",{"type":122,"code":226},"curl https:\u002F\u002Fopenrouter.ai\u002Fapi\u002Fv1\u002Fchat\u002Fcompletions \\\n  -H \"Authorization: Bearer $OPENROUTER_API_KEY\" \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\n    \"models\": [\"~openai\u002Fgpt-sol-latest\", \"~anthropic\u002Fclaude-sonnet-latest\"],\n    \"messages\": [\n      { \"role\": \"user\", \"content\": \"Summarise this stack trace in one sentence.\" }\n    ],\n    \"provider\": {\n      \"sort\": \"price\",\n      \"require_parameters\": true\n    }\n  }'",{"type":103,"content":228},[229,230,233],"The answer comes back in the OpenAI shape, and the ",{"tag":122,"children":231},[232],"model"," field names the endpoint that actually served it. That is also the model the request is billed at, whether or not it was first in the list. An existing OpenAI client needs only a different base URL; the documentation shows the same swap in Python and TypeScript.",{"type":235,"variant":236,"title":237,"body":238},"callout","tip","Aliases versus pins",[239],[240,241,244,245,248],"Slugs beginning with ",{"tag":122,"children":242},[243],"~"," resolve to the newest model in a family, and the documentation’s own examples use ",{"tag":122,"children":246},[247],"~openai\u002Fgpt-sol-latest"," so deployments pick up releases without a redeploy. The mechanism cuts both ways: when an invoice, a benchmark or a reproduction has to name an exact model, use the full identifier and expect a 404 once it is retired.",{"type":110,"level":111,"id":87,"text":88},{"type":103,"content":251},[252],"There is no subscription, no minimum and no lock-in. Credits are bought upfront by card, AliPay or USDC, and each request deducts the provider list price from the balance; the platform fee, 5.5% on Standard and 8% on Business, is charged on the purchase and never on the request.",{"type":116,"ordered":117,"items":254},[255,257,259,268],[256],"Free: 25+ free models, 4 providers, activity logs with export, 50 requests per day on free models. No auto-routing, no budgets, no prompt caching.",[258],"Standard: the full 500+ models and 80+ providers at a 5.5% fee on credit purchases, with auto-routing, budgets, prompt caching, management API keys and five workspaces.",[260,261,216,264,267],"Business: the same catalogue at 8%, EU or US in-region routing through ",{"tag":122,"children":262},[263],"eu.openrouter.ai",{"tag":122,"children":265},[266],"us.openrouter.ai",", 1,000 workspaces and workload identity federation.",[269],"Enterprise: contracted fees, invoicing, SSO and SCIM, contractual SLAs, and $200,000 of BYOK list-price inference per month before a 5% charge applies.",{"type":235,"variant":271,"title":272,"body":273},"note","Where the percentage lands",[274],[275],"Because the fee is fixed at the purchase, overhead on inference is a flat percentage regardless of which model answers: $10,000 of list-price traffic on Standard costs $550 a month. Paid plans also buy routing features rather than throughput — OpenRouter sets no plan-based rate limit, so a limit hit on a paid model normally comes from the provider.",{"type":110,"level":111,"id":90,"text":91},{"type":103,"content":278},[279],"For teams, the parts that matter are not in the request body. Workspaces, budgets and guardrails decide who may call what with which key, and they live in the platform rather than in your code, which is the point: policy changes stop being deployments.",{"type":116,"ordered":117,"items":281},[282,284,286,288],[283],"Workspaces separate environments: five on Free and Standard, 1,000 on Business, each with its own keys and limits, plus a credit cap per key.",[285],"Guardrails attach to keys or members and cover spend, model access, prompt-injection patterns and sensitive-information rules.",[287],"Management API keys create, list and delete keys programmatically, so key lifecycle belongs in CI rather than in a dashboard habit.",[289],"Broadcast ships traces to Datadog, Grafana Cloud, Langfuse, OpenTelemetry, Sentry, Snowflake or any HTTP endpoint, which is what makes spend joinable with the rest of the stack.",{"type":103,"content":291},[292,293,295],"Zero data retention runs on every plan, account-wide or per request with ",{"tag":122,"children":294},[215],". Regional residency is the feature that is not free: keeping prompts and completions inside the EU or the US starts at Business, which makes the 8% plan the baseline rather than the upgrade for regulated workloads.",{"type":110,"level":111,"id":93,"text":94},{"type":103,"content":298},[299,300,302],"The weaknesses are structural. Every request depends on a third party’s availability, and the routing decision stays invisible unless the metadata header is enabled. Parameter support differs between providers of the same model, so a call that works through one endpoint can be quietly degraded through another unless ",{"tag":122,"children":301},[210]," is set. Provider price changes flow straight through — the FAQ states that you will be charged at the new rate — and the free tier is rate-limited rather than generous. There is also ownership: Bloomberg and the Wall Street Journal reported in August 2026 that Stripe had agreed to buy the company, a useful reminder that the layer inside the request path has an owner.",{"type":304,"head":305,"rows":313},"table",[306,308,309,311],[307],"",[24],[310],"LiteLLM",[312],"First-party APIs",[314,323,332,341,350],[315,317,319,321],[316],"Deployment",[318],"Hosted, no self-hosted edition",[320],"Self-hosted proxy you operate",[322],"Your code per provider",[324,326,328,330],[325],"Fee",[327],"5.5% on credits (Standard)",[329],"Free under MIT, paid tiers by quote",[331],"None",[333,335,337,339],[334],"Catalogue",[336],"500+ models, 80+ providers",[338],"140+ providers, ~1,900 models",[340],"One vendor per integration",[342,344,346,348],[343],"Fallback",[345],"Per request, cross-provider and cross-model",[347],"Configured in the router",[349],"Written by hand",[351,353,355,357],[352],"Best fit",[354],"Many models, one bill, no operations",[356],"Platform team owning keys and budgets",[358],"One model on a hot path",{"type":103,"content":360},[361],"The comparison that matters is custody rather than features. OpenRouter removes the operations and adds a percentage plus a dependency; LiteLLM keeps the keys and the workload inside your infrastructure; first-party APIs give the shortest path and the least cover when a model is unavailable. For a product that changes models as often as others change a pricing page, that cover is worth more than the fee.",{"type":235,"variant":363,"title":364,"body":365},"warn","Pin it before it pins you",[366],[367,368,160,370,372,373,375,376,378],"Default load balancing can move the next request to a different provider, which changes latency, quantisation and sometimes output. Set ",{"tag":122,"children":369},[159],{"tag":122,"children":371},[140]," explicitly, keep ",{"tag":122,"children":374},[210]," on for structured outputs, and decide on purpose whether ",{"tag":122,"children":377},[205]," may change the model family mid-request — by default an error moves the call to the next model in the list.",{"type":110,"level":111,"id":96,"text":97},{"type":103,"content":381},[382],"OpenRouter answers how to give many agents, services and teams access to a fast-changing model list without operating a gateway. It does not answer how to own the request path.",{"type":116,"ordered":384,"items":385},true,[386,391,393,395,397],[387,388,390],"Adopt it when the model list changes faster than the integration: one endpoint and a ",{"tag":122,"children":389},[134]," array turn a vendor evaluation into a configuration value.",[392],"Adopt it for agent and coding-agent workloads, where the bill spans many models and one activity log with export beats four provider dashboards.",[394],"Use BYOK when contracts or data policy require a direct relationship with the provider; routing and analytics are kept, and the first $25,000 of list-price inference per month is fee-free.",[396],"Skip it on latency-critical paths that cannot pin a provider: availability routing is the product, and a request that may change endpoint is a tail latency nobody controls.",[398],"Skip it when one provider and one model is the whole product: a first-party SDK, direct billing and no intermediary percentage is less machinery for the same call.",{"type":400,"content":401},"quote",[402,403,404,408],"“Inference is billed at the provider’s list price on every plan.”"," ",{"tag":405,"href":32,"children":406},"a",[407],"the pricing page"," The fee is small and legible; the dependency is the part that has to be priced in, and it is not listed anywhere.",{"type":110,"level":111,"id":99,"text":100},{"type":116,"ordered":384,"items":411},[412,415,418,421,424,427],[413],{"tag":405,"href":29,"children":414},[28],[416],{"tag":405,"href":32,"children":417},[31],[419],{"tag":405,"href":35,"children":420},[34],[422],{"tag":405,"href":38,"children":423},[37],[425],{"tag":405,"href":41,"children":426},[40],[428],{"tag":405,"href":25,"children":429},[43],[431,480,559,620],{"slug":432,"published":433,"minutes":434,"category":7,"tags":435,"keywords":441,"about":448,"sources":452,"cover":471,"og":472,"expertise":46,"locales":473,"lang":48,"title":474,"description":475,"coverAlt":476,"url":477,"pricing":478,"kind":479},"ollama","2026-09-29",11,[436,437,438,439,440],"Local inference","Open models","llama.cpp","GGUF","Model serving",[432,442,443,444,445,446,447],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[449],{"name":450,"url":451},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[453,456,459,462,465,468],{"title":454,"url":455},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":457,"url":458},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":460,"url":461},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":463,"url":464},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":466,"url":467},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":469,"url":470},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[48,49,50],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":481,"published":482,"minutes":434,"category":7,"tags":483,"keywords":488,"about":495,"sources":502,"cover":552,"og":553,"expertise":46,"locales":554,"lang":48,"title":555,"description":556,"coverAlt":557,"url":498,"pricing":558,"kind":9},"portkey","2026-09-28",[9,484,485,486,487],"Guardrails","Routing","Observability","Cost control",[489,490,16,491,492,493,494],"portkey ai gateway","portkey vs litellm","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[496,499],{"name":497,"url":498},"Portkey","https:\u002F\u002Fportkey.ai",{"name":500,"url":501},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[503,506,509,512,515,518,521,524,527,530,533,536,539,542,545,548,551],{"title":504,"url":505},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":507,"url":508},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":510,"url":511},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":513,"url":514},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":516,"url":517},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":519,"url":520},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":522,"url":523},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":525,"url":526},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":528,"url":529},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":531,"url":532},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":534,"url":535},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":537,"url":538},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":540,"url":541},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":543,"url":544},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":546,"url":547},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":549,"url":550},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":31,"url":32},"\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[48,49,50],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":560,"published":561,"minutes":6,"category":7,"tags":562,"keywords":568,"about":576,"sources":588,"cover":613,"og":614,"expertise":46,"locales":615,"lang":48,"title":616,"description":617,"coverAlt":618,"url":579,"pricing":619,"kind":563},"langfuse","2026-08-13",[563,564,565,566,567],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[560,569,570,571,572,573,574,575],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[577,580,582,585],{"name":578,"url":579},"Langfuse","https:\u002F\u002Flangfuse.com",{"name":565,"url":581},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":583,"url":584},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":586,"url":587},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[589,592,595,598,601,604,607,610],{"title":590,"url":591},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":593,"url":594},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":596,"url":597},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":599,"url":600},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":602,"url":603},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":605,"url":606},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":608,"url":609},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":611,"url":612},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[48,49,50],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":621,"published":622,"minutes":6,"category":7,"tags":623,"keywords":627,"about":634,"sources":642,"cover":666,"og":667,"expertise":46,"locales":668,"lang":48,"title":669,"description":670,"coverAlt":671,"url":637,"pricing":672,"kind":673},"braintrust","2026-07-07",[624,486,625,626,564],"Evaluations","LLM-as-a-judge","CI gates",[621,628,629,630,631,632,633],"braintrust pricing","braintrust eval","autoevals library","braintrust vs langfuse","llm evaluation platform","eval driven development",[635,638,641],{"name":636,"url":637},"Braintrust","https:\u002F\u002Fwww.braintrust.dev",{"name":639,"url":640},"Continuous integration","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FContinuous_integration",{"name":586,"url":587},[643,646,649,652,655,658,661,663],{"title":644,"url":645},"Braintrust pricing: plans, credits and usage rates","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing",{"title":647,"url":648},"Braintrust documentation: plans and limits","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fplans-and-limits",{"title":650,"url":651},"Braintrust documentation: evaluation quickstart","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fevaluation-quickstart",{"title":653,"url":654},"Braintrust documentation: get started","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs",{"title":656,"url":657},"GitHub: Braintrust organisation repositories","https:\u002F\u002Fgithub.com\u002Forgs\u002Fbraintrustdata\u002Frepositories",{"title":659,"url":660},"Arize Phoenix documentation: self-hosting","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix\u002Fself-hosting",{"title":662,"url":597},"Langfuse pricing: cloud plans and billable units",{"title":664,"url":665},"LangChain pricing: LangSmith plans","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fog.jpg",[48,49,50],"Braintrust review: eval-first observability with a hard meter","Braintrust turns production traces into datasets and gated experiments. What Starter and Pro really include, which parts are open source, and where Phoenix, Langfuse and LangSmith win.","A loop from instrumented application logs into a dataset, an experiment with scorers, and a comparison that gates the pull request before the change returns to the application.","Free · from $249 per month","Evaluation platform",1791383548847]