[{"data":1,"prerenderedAt":559},["ShallowReactive",2],{"tool-litellm-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":29,"cover":51,"og":52,"expertise":53,"locales":54,"lang":55,"title":58,"description":59,"coverAlt":60,"url":61,"pricing":62,"kind":9,"metaTitle":63,"takeaways":64,"faq":70,"toc":83,"blocks":108,"others":331},"litellm","2026-05-20",10,"llmops",[9,10,11,12,13],"LLM gateway","OpenAI-compatible","Cost control","Routing","MCP",[4,15,16,17,18,19,20,21],"litellm vs portkey","llm gateway comparison","openai compatible proxy","litellm proxy config yaml","llm spend tracking per team","self-hosted llm gateway","llm budget enforcement",[23,26],{"name":24,"url":25},"LiteLLM","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLiteLLM",{"name":27,"url":28},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",[30,33,36,39,42,45,48],{"title":31,"url":32},"LiteLLM documentation: getting started","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002F",{"title":34,"url":35},"LiteLLM documentation: virtual keys","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002Fproxy\u002Fvirtual_keys",{"title":37,"url":38},"LiteLLM documentation: router and load balancing","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002Frouting",{"title":40,"url":41},"LiteLLM documentation: benchmarks","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002Fbenchmarks",{"title":43,"url":44},"LiteLLM documentation: supported endpoints","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002Fsupported_endpoints",{"title":46,"url":47},"LiteLLM documentation: Enterprise","https:\u002F\u002Fdocs.litellm.ai\u002Fdocs\u002Fenterprise",{"title":49,"url":50},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Flitellm\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flitellm\u002Fog.jpg","ai-engineer",[55,56,57],"en","de","hu","LiteLLM: the OpenAI-compatible gateway most platform teams end up running","LiteLLM is the MIT-licensed OpenAI-compatible gateway most platform teams put in front of their providers. What it does well, what it costs and where it breaks.","Request path through the LiteLLM gateway: application, key check, router, provider, response, with Postgres and Redis underneath.","https:\u002F\u002Fwww.litellm.ai","MIT · Enterprise from $20 per seat","LiteLLM review: the gateway in the middle · Balázs Csorba",[65,66,67,68,69],"LiteLLM is an MIT-licensed Python SDK and proxy that puts 140+ providers behind one OpenAI-shaped endpoint, with virtual keys, per-team budgets and cross-provider fallbacks.","The value is organisational rather than technical: changing model becomes a config edit, and every request produces a spend row.","Overhead is measurable and published: the docs report 8ms p95 at 1,000 requests per second, and are unusually explicit that latency only means something next to the number of requests in flight.","The free tier is the whole gateway. SSO beyond five users, audit logs, key rotation and the built-in moderation callbacks sit behind a licence quoted by annual request capacity, with no public list price.","Release cadence is roughly one minor line a week against a four-line support window, so the upgrade is a recurring operational task rather than an event.",[71,74,77,80],{"q":72,"a":73},"Is LiteLLM just a wrapper around the OpenAI SDK?","No. The SDK normalises requests, responses and exceptions across the supported providers, but the proxy adds virtual keys, spend tracking, budgets, routing, cooldowns, fallbacks and caching on top. The usual path is to start with the SDK inside one service and move to the proxy once more than one team needs access.",{"q":75,"a":76},"How much latency does the gateway add?","The benchmark page reports 8ms p95 at 1,000 requests per second on four instances of 4 vCPU and 8 GB. That figure includes the client: the docs point out that with no think time a closed-loop client holds 1,000 requests in flight instead of about 130, which by Little's Law reports roughly eight times the latency at the same throughput.",{"q":78,"a":79},"Does the free version really enforce budgets?","Yes. Virtual keys, teams, spend tracking, budgets, rate limits and LLM fallbacks are all in the open-source tier. What the licence gates is identity and governance: SSO beyond five users, SCIM, audit logs, secret-manager write-back, virtual key rotation and the built-in moderation callbacks.",{"q":81,"a":82},"How is Enterprise priced?","There is no public list price. The pricing page says the licence is sized to annual gateway request capacity, deployment architecture and support needs, never per token, and every quote comes from sales. A 30-day trial key is issued without a credit card, and procurement is also available through AWS Marketplace and resellers.",[84,87,90,93,96,99,102,105],{"id":85,"title":86},"what-it-is","What LiteLLM is",{"id":88,"title":89},"how-it-works","How it works",{"id":91,"title":92},"getting-started","Getting started",{"id":94,"title":95},"keys-and-budgets","Keys, budgets and spend",{"id":97,"title":98},"pricing","Pricing and licence boundaries",{"id":100,"title":101},"where-it-shingles","Where it shingles",{"id":103,"title":104},"verdict","Verdict",{"id":106,"title":107},"sources","Sources",[109,113,116,119,122,140,141,144,153,156,160,163,166,178,179,182,184,191,213,214,217,219,222,223,226,229,230,233,280,283,284,287,300,306,307],{"type":110,"content":111},"paragraph",[112],"LiteLLM is the open-source LLM gateway that most platform teams end up running whether or not they meant to. It is two products in one package: a Python SDK that normalises the major providers onto the OpenAI request and response shape, and a proxy server that puts virtual keys, budgets, routing and spend tracking in front of them. On the documentation and the published numbers it is the most complete piece of plumbing in this category, and heavy enough that it deserves a deliberate decision rather than an accidental one.",{"type":110,"content":114},[115],"It sits one hop in front of the model providers, not in front of the application data. Retrieval, prompt management and agent orchestration live elsewhere; an OpenAI client, a LangGraph workflow and a LiteLLM proxy coexist without complaint. What the gateway owns is the request itself: who is allowed to make it, which deployment serves it, what it costs and who is accountable for the bill.",{"type":117,"level":118,"id":85,"text":86},"heading",2,{"type":110,"content":120},[121],"The name covers two runnable things that share a config file and a model price table. The SDK is a library you import. The proxy is a server that speaks the OpenAI wire format, so any client that already works against OpenAI can be redirected by changing its base URL and its key.",{"type":123,"ordered":124,"items":125},"list",false,[126,128,130,132,134,136,138],[127],"MIT-licensed, with an Enterprise tier that adds identity and governance features rather than capacity.",[129],"140+ provider integrations and roughly 1,900 models, tracked in a public price and context-window file that ships with the package.",[131],"OpenAI-compatible endpoints including chat completions, the Anthropic messages route, the responses API, embeddings, batches and realtime.",[133],"Virtual keys with per-key, per-team and per-user budgets, rate limits and model allowlists, all backed by Postgres.",[135],"A router with several deployment-selection strategies, per-deployment cooldowns, weighted failover, cross-model fallbacks and session affinity.",[137],"An MCP gateway and an A2A agent gateway on the same endpoint and the same key policy, so agents get governed tool access rather than a separate one.",[139],"Callbacks into Langfuse, MLflow, Helicone, LangSmith and OpenTelemetry, plus Prometheus metrics and an admin UI.",{"type":117,"level":118,"id":88,"text":89},{"type":110,"content":142},[143],"Every request through the proxy follows the same order of operations: authenticate the key, check the budget and rate limits, count tokens, pick a deployment, call the provider, translate the response, record the cost. Each step can fail on its own, and the router's behaviour on failure is where most of the operational learning happens.",{"type":145,"attrs":146,"inner":150,"caption":151},"diagram",{"viewBox":147,"role":148,"aria-labelledby":149},"0 0 720 320","img","d1-llm-t d1-llm-d","\u003Ctitle id=\"d1-llm-t\">LiteLLM request path\u003C\u002Ftitle>\u003Cdesc id=\"d1-llm-d\">An application sends a chat completion request to the LiteLLM gateway. The gateway checks the virtual key and the budget, the router picks a deployment, the provider answers, and the same OpenAI-shaped JSON goes back. Postgres, Redis and observability callbacks sit underneath the request path.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">LiteLLM request path\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">openai-compatible\u003C\u002Ftext>\u003Crect x=\"20\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"78\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">App\u003C\u002Ftext>\u003Ctext x=\"78\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">OpenAI client\u003C\u002Ftext>\u003Cpath d=\"M136 102 H148\" class=\"d-line\" \u002F>\u003Cpath d=\"M156 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"156\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"214\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Gateway\u003C\u002Ftext>\u003Ctext x=\"214\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">key, budget\u003C\u002Ftext>\u003Cpath d=\"M272 102 H284\" class=\"d-line\" \u002F>\u003Cpath d=\"M292 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"292\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"350\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Router\u003C\u002Ftext>\u003Ctext x=\"350\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">pick deployment\u003C\u002Ftext>\u003Cpath d=\"M408 102 H420\" class=\"d-line\" \u002F>\u003Cpath d=\"M428 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"428\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"486\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Provider\u003C\u002Ftext>\u003Ctext x=\"486\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">OpenAI, Bedrock\u003C\u002Ftext>\u003Cpath d=\"M544 102 H556\" class=\"d-line\" \u002F>\u003Cpath d=\"M564 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"564\" y=\"70\" width=\"116\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"622\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Response\u003C\u002Ftext>\u003Ctext x=\"622\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">same JSON shape\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"172\" class=\"d-label\">STATE AND OBSERVABILITY\u003C\u002Ftext>\u003Cpath d=\"M214 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M214 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M350 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M350 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M601 216 V146\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M601 138 l-5 9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"134\" y=\"216\" width=\"160\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"214\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Postgres\u003C\u002Ftext>\u003Ctext x=\"214\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">keys, spend rows\u003C\u002Ftext>\u003Crect x=\"270\" y=\"216\" width=\"160\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"350\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Redis\u003C\u002Ftext>\u003Ctext x=\"350\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">cooldowns, cache\u003C\u002Ftext>\u003Crect x=\"502\" y=\"216\" width=\"198\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"601\" y=\"244\" text-anchor=\"middle\" class=\"d-text\">Callbacks\u003C\u002Ftext>\u003Ctext x=\"601\" y=\"265\" text-anchor=\"middle\" class=\"d-small\">Langfuse, OTEL\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"306\" text-anchor=\"middle\" class=\"d-label\">The gateway never holds the provider keys\u003C\u002Ftext>",[152],"One hop in front of every provider. Postgres and Redis are what turn a request translator into a control point, and they are also what has to be sized and backed up.",{"type":110,"content":154},[155],"Model names are the abstraction that makes this work. The config file lists deployments under a shared alias, several deployments may carry the same alias, and a request for that alias is answered by whichever deployment the routing strategy picks. Move the alias and every caller changes provider without a code change. The same table holds the fallbacks: a failed deployment is either re-picked from its peers or escalated to a named fallback model.",{"type":117,"level":157,"id":158,"text":159},3,"overhead","Measured overhead",{"type":110,"content":161},[162],"The benchmark page reports 8ms p95 at 1,000 requests per second across four instances of 4 vCPU and 8 GB, and for the realtime endpoint 59ms median, 67ms p95 and 99ms p99 at 1,207 RPS. The same page publishes a head-to-head against Portkey at roughly 1,170 RPS on the same hardware, where LiteLLM reports p95 150ms and p99 240ms against Portkey's 230ms and 500ms.",{"type":110,"content":164},[165],"The more useful part is the caveat the docs volunteer. The load generator sleeps 0.5 to 1 second between requests, so about 130 requests are in flight at any instant; a closed-loop client with no think time holds 1,000 and therefore reports roughly eight times the latency at the same throughput. A published overhead number is meaningless without the in-flight depth next to it, which is more disclosure than most vendor benchmarks carry.",{"type":167,"variant":168,"title":169,"body":170},"callout","note","Measure it yourself",[171],[172,173,177],"The docs ship a ",{"tag":174,"children":175},"code",[176],"network_mock"," mode that intercepts outbound requests at the HTTP transport layer and returns canned responses, plus a public benchmark script. Running it against your own traffic takes minutes and produces the only overhead figure worth planning around.",{"type":117,"level":118,"id":91,"text":92},{"type":110,"content":180},[181],"The smallest useful deployment is a config file with two deployments behind one alias, a master key and a Postgres connection. Caching, guardrails, callbacks and the admin UI all come later.",{"type":174,"code":183},"# config.yaml: two deployments behind one alias, plus a managed fallback\nmodel_list:\n  - model_name: chat-fast\n    litellm_params:\n      model: openai\u002Fgpt-5.6-luna\n      api_key: os.environ\u002FOPENAI_API_KEY\n      rpm: 900\n  - model_name: chat-fast\n    litellm_params:\n      model: anthropic\u002Fclaude-sonnet-5\n      api_key: os.environ\u002FANTHROPIC_API_KEY\n      rpm: 60\n  - model_name: chat-cheap\n    litellm_params:\n      model: openai\u002Fgpt-5.6-luna\n      api_key: os.environ\u002FOPENAI_API_KEY\n\nrouter_settings:\n  routing_strategy: simple-shuffle\n  num_retries: 2\n  fallbacks:\n    - chat-fast: [chat-cheap]\n\ngeneral_settings:\n  master_key: os.environ\u002FLITELLM_MASTER_KEY\n  database_url: os.environ\u002FDATABASE_URL\n\n# docker run -p 4000:4000 -v $(pwd)\u002Fconfig.yaml:\u002Fapp\u002Fconfig.yaml \\\n#   -e DATABASE_URL -e LITELLM_MASTER_KEY \\\n#   docker.litellm.ai\u002Fberriai\u002Flitellm:latest --config \u002Fapp\u002Fconfig.yaml",{"type":110,"content":185},[186,187,190],"The client change is two lines: point the OpenAI SDK at the proxy and pass a virtual key instead of the provider key. Every response carries an ",{"tag":174,"children":188},[189],"x-litellm-overhead-duration-ms"," header reporting what the gateway itself cost, which is the number to alert on when latency regresses and the upstream did not change.",{"type":167,"variant":192,"title":193,"body":194},"warn","Retries compound, and concurrency caps now reject",[195],[196,197,200,201,204,205,208,209,212],"The docs are explicit that ",{"tag":174,"children":198},[199],"num_retries"," is LiteLLM's own loop and that the provider SDK's ",{"tag":174,"children":202},[203],"max_retries"," is pinned to zero on routed calls, so a deployment value cannot be applied twice. Separately, a deployment whose ",{"tag":174,"children":206},[207],"max_parallel_requests"," or inferred ",{"tag":174,"children":210},[211],"rpm"," cap is saturated now returns 429 immediately instead of queueing. Earlier versions waited for a free slot, so code that relied on the queue needs the cap raised rather than the timeout lengthened.",{"type":117,"level":118,"id":94,"text":95},{"type":110,"content":215},[216],"The key model is where LiteLLM earns its keep. A virtual key carries a model allowlist, a maximum budget with a reset window, TPM and RPM limits, a parallel request cap and an owner. Keys inherit from their owner, which is convenient and also the sharpest edge in the design: a key minted by an admin without a user id inherits nothing at all, and a key owned by a proxy admin can reach the admin management routes because management access follows the owner's role rather than the key's own permissions.",{"type":174,"code":218},"# mint a key for one team, capped, rate-limited and attributed\ncurl 'http:\u002F\u002Flocalhost:4000\u002Fkey\u002Fgenerate' \\\n  -H \"Authorization: Bearer $LITELLM_MASTER_KEY\" \\\n  -H 'Content-Type: application\u002Fjson' \\\n  -d '{\n    \"models\": [\"chat-fast\"],\n    \"max_budget\": 40,\n    \"budget_duration\": \"30d\",\n    \"rpm_limit\": 600,\n    \"tpm_limit\": 400000,\n    \"metadata\": {\"team\": \"platform\", \"env\": \"prod\"}\n  }'\n\n# what has this key spent, and against which cap\ncurl 'http:\u002F\u002Flocalhost:4000\u002Fkey\u002Finfo?key=sk-...' \\\n  -H \"Authorization: Bearer $LITELLM_MASTER_KEY\"",{"type":110,"content":220},[221],"Spend is computed from the same public price file the SDK ships, written on every completion, embedding and image call, and queryable per key, user and team. Those numbers are therefore estimates derived from a price table that can lag a provider's own billing: close enough to charge a team back, not close enough to reconcile an invoice. When a cost figure matters for a contract, take it from the provider invoice.",{"type":117,"level":118,"id":97,"text":98},{"type":110,"content":224},[225],"The gateway is free to self-host forever under MIT, with no per-seat and no per-token charge. The pricing page lists the open-source tier as 140+ provider integrations, virtual keys, users and teams, spend tracking, budgets and rate limits, LLM fallbacks, request and response logging, Prometheus metrics and guardrails. Procurement runs directly or through AWS Marketplace and resellers.",{"type":110,"content":227},[228],"The licence buys control, not throughput. SSO is included up to five users and needs a licence beyond that; SCIM, audit logs, secret-manager write-back, IP allowlists, the multi-region control plane, virtual key rotation and the built-in moderation callbacks all sit behind it, and the docs name them individually so nobody discovers the boundary after rollout. Enterprise pricing is quoted against annual request capacity and deployment shape, never per token, and no list price is published, so any real cost comparison starts with a sales conversation.",{"type":117,"level":118,"id":100,"text":101},{"type":110,"content":231},[232],"The weaknesses are operational rather than functional. The proxy needs Postgres and, once it runs on more than one instance, Redis; that is a stateful system with schema migrations and connection arithmetic, and the docs carry dedicated sizing pages precisely because it is where deployments fail. The Python SDK pulls a large dependency tree into a service for what is mostly request formatting. And the March 2026 supply-chain incident, in which trojanised LiteLLM releases reached PyPI, is a reminder that this is a package with very high install counts handling production keys.",{"type":234,"head":235,"rows":243},"table",[236,238,239,241],[237],"",[24],[240],"Portkey",[242],"Bifrost",[244,253,262,271],[245,247,249,251],[246],"Deployment",[248],"Self-hosted, air-gapped supported",[250],"Managed cloud and self-hosted",[252],"Self-hosted",[254,256,258,260],[255],"Cost basis",[257],"Free OSS, licence by capacity",[259],"Per-month subscription",[261],"No published list price",[263,265,267,269],[264],"Vendor p99 overhead",[266],"0.66 ms",[268],"2.29 ms",[270],"4.54 ms",[272,274,276,278],[273],"Best fit",[275],"Platform team owning access",[277],"Guardrails first",[279],"Minimal hop overhead",{"type":110,"content":281},[282],"Those overhead figures come from LiteLLM's own benchmark, run with every gateway pointed at the same deterministic mock upstream on identical hardware, which makes the comparison fair and the winner inevitable. For anything that is not an LLM gateway, the Kubernetes-native option deserves a look: Envoy AI Gateway exposes the OpenAI routing surface without adding a Python runtime to the request path, which is the structural difference that matters more than any of the numbers above.",{"type":117,"level":118,"id":103,"text":104},{"type":110,"content":285},[286],"LiteLLM is the right default for a platform team that must give many teams access to several providers and needs three questions answered in production: who spent this, which deployment served it, and what happens when that provider degrades. It answers all three better than the alternatives, and it does so without taking custody of the provider keys.",{"type":123,"ordered":288,"items":289},true,[290,292,294,296,298],[291],"Adopt it when more than one team calls models through more than one provider. Budgets, spend attribution and cross-provider fallback are what justify the stateful deployment.",[293],"Adopt it when procurement consolidation is the goal. Moving from one model to another becomes a config edit rather than a security review cycle.",[295],"Take the SDK alone first if there is a single service. It normalises errors and responses for a fraction of the operational cost and leaves the door open to the proxy later.",[297],"Skip it if one team and one provider is the whole story. A provider SDK and a spend spreadsheet are less machinery for the same result.",[299],"Skip it if the last millisecond matters on the hot path. The gateway is a Python process in the request path, and the Rust rewrite is still labelled beta.",{"type":167,"variant":301,"title":302,"body":303},"tip","Before committing",[304],[305],"Pin a minor line, take its patches, and move up before it drops out of the four-line support window. Read the database sizing page before the first scale-out rather than after the connection errors start. And if the provider SDK is already installed everywhere, pin the exact version in a lockfile rather than trusting the resolver.",{"type":117,"level":118,"id":106,"text":107},{"type":123,"ordered":288,"items":308},[309,313,316,319,322,325,328],[310],{"tag":311,"href":32,"children":312},"a",[31],[314],{"tag":311,"href":35,"children":315},[34],[317],{"tag":311,"href":38,"children":318},[37],[320],{"tag":311,"href":41,"children":321},[40],[323],{"tag":311,"href":44,"children":324},[43],[326],{"tag":311,"href":47,"children":327},[46],[329],{"tag":311,"href":50,"children":330},[49],[332,381,456,517],{"slug":333,"published":334,"minutes":335,"category":7,"tags":336,"keywords":342,"about":349,"sources":353,"cover":372,"og":373,"expertise":53,"locales":374,"lang":55,"title":375,"description":376,"coverAlt":377,"url":378,"pricing":379,"kind":380},"ollama","2026-09-29",11,[337,338,339,340,341],"Local inference","Open models","llama.cpp","GGUF","Model serving",[333,343,344,345,346,347,348],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[350],{"name":351,"url":352},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[354,357,360,363,366,369],{"title":355,"url":356},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":358,"url":359},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":361,"url":362},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":364,"url":365},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":367,"url":368},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":370,"url":371},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[55,56,57],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":382,"published":383,"minutes":335,"category":7,"tags":384,"keywords":387,"about":393,"sources":399,"cover":449,"og":450,"expertise":53,"locales":451,"lang":55,"title":452,"description":453,"coverAlt":454,"url":395,"pricing":455,"kind":9},"portkey","2026-09-28",[9,385,12,386,11],"Guardrails","Observability",[388,389,16,390,391,20,392],"portkey ai gateway","portkey vs litellm","llm gateway latency overhead","llm guardrails gateway","portkey pricing",[394,396],{"name":240,"url":395},"https:\u002F\u002Fportkey.ai",{"name":397,"url":398},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[400,403,406,409,412,415,418,421,424,427,430,433,436,439,442,445,446],{"title":401,"url":402},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":404,"url":405},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":407,"url":408},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":410,"url":411},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":413,"url":414},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":416,"url":417},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":419,"url":420},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":422,"url":423},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":425,"url":426},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":428,"url":429},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":431,"url":432},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":434,"url":435},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":437,"url":438},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":440,"url":441},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":443,"url":444},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":49,"url":50},{"title":447,"url":448},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[55,56,57],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":457,"published":458,"minutes":6,"category":7,"tags":459,"keywords":465,"about":473,"sources":485,"cover":510,"og":511,"expertise":53,"locales":512,"lang":55,"title":513,"description":514,"coverAlt":515,"url":476,"pricing":516,"kind":460},"langfuse","2026-08-13",[460,461,462,463,464],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[457,466,467,468,469,470,471,472],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[474,477,479,482],{"name":475,"url":476},"Langfuse","https:\u002F\u002Flangfuse.com",{"name":462,"url":478},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":480,"url":481},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":483,"url":484},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[486,489,492,495,498,501,504,507],{"title":487,"url":488},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":490,"url":491},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":493,"url":494},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":496,"url":497},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":499,"url":500},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":502,"url":503},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":505,"url":506},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":508,"url":509},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[55,56,57],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":518,"published":519,"minutes":6,"category":7,"tags":520,"keywords":524,"about":531,"sources":535,"cover":551,"og":552,"expertise":53,"locales":553,"lang":55,"title":554,"description":555,"coverAlt":556,"url":557,"pricing":558,"kind":9},"openrouter","2026-07-23",[9,521,522,10,523],"Model routing","Fallbacks","Pay per token",[518,525,16,526,527,528,529,530],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[532],{"name":533,"url":534},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[536,539,540,543,546,549],{"title":537,"url":538},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":447,"url":448},{"title":541,"url":542},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":544,"url":545},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":547,"url":548},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":550,"url":534},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[55,56,57],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",1791383549078]