[{"data":1,"prerenderedAt":755},["ShallowReactive",2],{"tool-vllm-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":21,"sources":25,"cover":53,"og":54,"expertise":55,"locales":56,"lang":57,"title":60,"description":61,"coverAlt":62,"url":63,"pricing":64,"kind":65,"metaTitle":66,"takeaways":67,"faq":73,"toc":89,"blocks":114,"others":519},"vllm","2026-06-30",11,"llmops",[9,10,11,12,13],"Self-hosted inference","OpenAI API","PagedAttention","GPU serving","LLM runtime",[4,15,16,17,18,19,20],"vllm vs sglang","vllm vs tensorrt-llm","self-hosted llm server","openai compatible api server","llm inference throughput","pagedattention",[22],{"name":23,"url":24},"vLLM","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FvLLM",[26,29,32,35,38,41,44,47,50],{"title":27,"url":28},"vLLM documentation","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Flatest\u002F",{"title":30,"url":31},"Optimization and tuning for the V1 engine","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Fconfiguration\u002Foptimization.html",{"title":33,"url":34},"Architecture overview and the V1 process layout","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Fdesign\u002Farch_overview.html",{"title":36,"url":37},"Metrics design","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Fdesign\u002Fmetrics.html",{"title":39,"url":40},"Automatic prefix caching and its documented limits","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Ffeatures\u002Fautomatic_prefix_caching.html",{"title":42,"url":43},"Online serving and the HTTP API surface","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Fserving\u002Fonline_serving.html",{"title":45,"url":46},"Inside vLLM: anatomy of a high-throughput inference system","https:\u002F\u002Fvllm.ai\u002Fblog\u002F2025-09-05-anatomy-of-vllm",{"title":48,"url":49},"vLLM: easy, fast and cheap LLM serving with PagedAttention","https:\u002F\u002Fblog.vllm.ai\u002F2023\u002F06\u002F20\u002Fvllm.html",{"title":51,"url":52},"Release history","https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\u002Freleases","\u002Fimages\u002Fblog\u002Fvllm\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fvllm\u002Fog.jpg","ai-engineer",[57,58,59],"en","de","hu","vLLM reviewed for self-hosted inference","vLLM turns a Hugging Face checkpoint into an OpenAI-compatible server. What PagedAttention and continuous batching buy, and what running it actually costs.","Schematic of the vLLM serving stack, from client requests through the scheduler to paged KV cache blocks","https:\u002F\u002Fdocs.vllm.ai","Apache-2.0","Inference server","vLLM reviewed for self-hosted inference · Balázs Csorba",[68,69,70,71,72],"vLLM is the default self-hosted serving engine for open-weight models on NVIDIA and AMD hardware, chosen for breadth of model, quantisation and API coverage rather than for being the fastest.","PagedAttention and continuous batching are the reason it exists; the durable claim is that paged allocation cut KV cache waste from 60 to 80 per cent down to under 4 per cent, not the 24x throughput multiple from 2023.","Prefix caching is enabled by default and only shortens prefill, so workloads with long repeated prompts gain the most and long generations with unique prompts gain nothing from it.","The production failure mode is preemption by recompute when the KV cache is undersized; watch KV cache usage and the cumulative preemption count rather than aggregate throughput.","A four-GPU deployment runs six processes and needs at least six physical CPU cores, which is the most common reason throughput lands below expectations.",[74,77,80,83,86],{"q":75,"a":76},"Is vLLM faster than SGLang or TensorRT-LLM?","Not categorically, and most published comparisons are not comparable. The 2023 vLLM launch post reported up to 24 times the throughput of Hugging Face Transformers and 2.2 to 3.5 times that of TGI, measured on LLaMA-7B and LLaMA-13B with request lengths sampled from ShareGPT. That predates competitors adopting paged caches, so benchmark your own request-length distribution rather than trusting any single multiple.",{"q":78,"a":79},"Do I need a GPU to run vLLM?","Linux with CUDA or ROCm is the documented target, with Python 3.10 to 3.13. Intel XPU and Google TPU are supported separately, and Apple Silicon is covered by a different, MLX-based project rather than by vLLM itself. CPU-only inference is llama.cpp territory, not vLLM's.",{"q":81,"a":82},"How much VRAM does vLLM need?","vLLM takes a configurable fraction of VRAM for weights and KV cache, then measures what is left by running a profiling forward pass at start-up and reports how many KV cache blocks fit. In practice the binding constraint is the model weights at your chosen quantisation; the remainder is KV cache, and how much that is decides your concurrency.",{"q":84,"a":85},"Can one vLLM server host several models?","Not natively. A server process hosts one model at a time, so multi-model deployments run one process per model behind a router, or use data parallelism to replicate the same model across several GPUs. LoRA adapters are the exception: they can be loaded and unloaded at runtime, and the documentation restricts that to local development.",{"q":87,"a":88},"How do I keep vLLM output reproducible across upgrades?","Pin the engine version, since minor releases land roughly every two weeks, and pass --generation-config vllm to stop the server applying the model's generation_config.json from Hugging Face, which otherwise overrides your sampling defaults silently. For deterministic runs, set temperature to zero and pin the model revision as well.",[90,93,96,99,102,105,108,111],{"id":91,"title":92},"what-it-is","What it actually is",{"id":94,"title":95},"how-it-works","How it works",{"id":97,"title":98},"getting-started","Getting a server up",{"id":100,"title":101},"performance","What the throughput claims are worth",{"id":103,"title":104},"running-in-production","Running it in production",{"id":106,"title":107},"where-it-stings","Where it stings",{"id":109,"title":110},"verdict","Verdict",{"id":112,"title":113},"sources","Sources",[115,119,122,125,128,149,150,153,162,165,166,169,172,199,201,216,217,220,286,289,307,308,312,323,357,360,363,369,372,383,403,404,415,461,464,465,468,483,488,489],{"type":116,"content":117},"paragraph",[118],"vLLM is an open-source inference engine that turns a Hugging Face checkpoint into an HTTP server that speaks the OpenAI API. It is the layer between a model and everything that calls it, and its entire job is to keep the GPU busy. The verdict up front: on NVIDIA or AMD hardware running open-weight models it is the default choice, because no other project covers this much of the surface with this little glue code. On a laptop, on a CPU-only host, or with one model and one user, it is far too much machinery.",{"type":116,"content":120},[121],"It competes in the serving layer, not the model layer. The alternatives are SGLang, which shares most of its design; NVIDIA's TensorRT-LLM, which trades breadth for peak numbers on NVIDIA parts; llama.cpp, which starts where vLLM gives up; and Hugging Face TGI, whose last release was v3.3.7 in December 2025. It is also the usual substitute for a hosted API, because the same OpenAI client code that talks to a provider can be pointed at a local process.",{"type":123,"level":124,"id":91,"text":92},"heading",2,{"type":116,"content":126},[127],"The facts that matter when picking a serving engine, all of them taken from the project's own documentation rather than from a vendor page:",{"type":129,"ordered":130,"items":131},"list",false,[132,137,139,141,143,145,147],[133,136],{"tag":134,"children":135},"strong",[64],", with no paid tier and no control plane hosted by the vendor.",[138],"Started at UC Berkeley's Sky Computing Lab; the documentation credits more than 2,000 contributors and calls it one of the most active open-source AI projects.",[140],"More than 200 model architectures on Hugging Face, spanning decoder-only, mixture-of-experts, hybrid attention, multimodal, embedding, rerank and reward models.",[142],"An OpenAI-compatible server plus Anthropic Messages and Cohere embed and rerank endpoints, and it also covers speech and structured output. One model per server process.",[144],"CUDA and ROCm as first-class targets, Intel XPU and Google TPU supported, with hardware plugins for Ascend NPUs, Gaudi, Spyre, Apple Silicon, MetaX and others.",[146],"Quantisation across FP8, NVFP4, MXFP4, INT8 and INT4, plus GPTQ, AWQ, GGUF and compressed-tensors checkpoints.",[148],"A minor release roughly every two weeks: v0.24.0 shipped on 29 June 2026, six weeks after v0.20.2 on 10 May.",{"type":123,"level":124,"id":94,"text":95},{"type":116,"content":151},[152],"Two ideas carry most of the throughput. PagedAttention stores the key and value cache in fixed-size blocks and maps them onto non-contiguous GPU memory through a block table, the way an operating system maps pages; the 2023 paper measured 60 to 80 per cent of KV cache wasted through fragmentation and over-reservation in earlier systems, against under 4 per cent for paged allocation. Continuous batching then admits new requests at every decode step instead of waiting for a batch to fill, which is where most of the latency improvement comes from.",{"type":154,"attrs":155,"inner":159,"caption":160},"diagram",{"viewBox":156,"role":157,"aria-labelledby":158},"0 0 720 300","img","vllm-flow-t vllm-flow-d","\u003Ctitle id=\"vllm-flow-t\">One request through the vLLM V1 engine\u003C\u002Ftitle>\u003Cdesc id=\"vllm-flow-d\">Clients call an API server process, which tokenises the prompt and forwards it over a socket to an engine core process. The engine core holds the scheduler and the KV cache manager, which hands blocks to one worker process per GPU. Token by token the workers write into paged KV cache blocks and streamed output travels back along the same path.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-vllm\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"20\" y=\"30\" class=\"d-title\">V1 REQUEST PATH\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"30\" text-anchor=\"end\" class=\"d-label\">one model, many tenants, one process per GPU\u003C\u002Ftext>\u003Crect x=\"20\" y=\"72\" width=\"120\" height=\"72\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"80\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">clients\u003C\u002Ftext>\u003Ctext x=\"80\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">OpenAI API\u003C\u002Ftext>\u003Cpath d=\"M142 108H166\" class=\"d-line\" marker-end=\"url(#ah-vllm)\" \u002F>\u003Crect x=\"170\" y=\"72\" width=\"140\" height=\"72\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"240\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">API server\u003C\u002Ftext>\u003Ctext x=\"240\" y=\"122\" text-anchor=\"middle\" class=\"d-small\">tokenise, stream\u003C\u002Ftext>\u003Cpath d=\"M312 108H346\" class=\"d-line\" marker-end=\"url(#ah-vllm)\" \u002F>\u003Ctext x=\"329\" y=\"98\" text-anchor=\"middle\" class=\"d-label\">ZMQ\u003C\u002Ftext>\u003Crect x=\"350\" y=\"72\" width=\"150\" height=\"72\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"425\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">engine core\u003C\u002Ftext>\u003Ctext x=\"425\" y=\"122\" text-anchor=\"middle\" class=\"d-small\">scheduler, KV manager\u003C\u002Ftext>\u003Cpath d=\"M502 108H536\" class=\"d-line\" marker-end=\"url(#ah-vllm)\" \u002F>\u003Crect x=\"540\" y=\"72\" width=\"160\" height=\"72\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"620\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">GPU workers\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"122\" text-anchor=\"middle\" class=\"d-small\">one process per GPU\u003C\u002Ftext>\u003Cpath d=\"M620 148V168H80V150\" class=\"d-line d-dash\" marker-end=\"url(#ah-vllm)\" \u002F>\u003Ctext x=\"350\" y=\"190\" text-anchor=\"middle\" class=\"d-small\">sampled token by token, streamed back through the same path\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"222\" class=\"d-label\">paged kv cache blocks, addressed through a block table\u003C\u002Ftext>\u003Crect x=\"20\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"40\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">0\u003C\u002Ftext>\u003Crect x=\"70\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"90\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">1\u003C\u002Ftext>\u003Crect x=\"120\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"140\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">2\u003C\u002Ftext>\u003Crect x=\"170\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"190\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">3\u003C\u002Ftext>\u003Crect x=\"220\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"240\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">4\u003C\u002Ftext>\u003Crect x=\"270\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-accent\" \u002F>\u003Ctext x=\"290\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">5\u003C\u002Ftext>\u003Crect x=\"320\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-accent\" \u002F>\u003Ctext x=\"340\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">6\u003C\u002Ftext>\u003Crect x=\"370\" y=\"234\" width=\"40\" height=\"34\" rx=\"6\" class=\"d-accent\" \u002F>\u003Ctext x=\"390\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">7\u003C\u002Ftext>\u003Ctext x=\"430\" y=\"250\" class=\"d-small\">a shared prefix maps many requests\u003C\u002Ftext>\u003Ctext x=\"430\" y=\"266\" class=\"d-small\">onto the same blocks\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"292\" class=\"d-small\">compute-bound prefill and memory-bound decode are scheduled into the same batch\u003C\u002Ftext>",[161],"vLLM V1 splits the work across processes: HTTP and tokenisation in the API server, scheduling and cache management in the engine core, one worker per GPU.",{"type":116,"content":163},[164],"The V1 engine, which replaced V0 during 2025, mixes prefill and decode in the same step and gives decode priority, so a long prompt no longer stalls streaming traffic. Chunked prefill is on by default. Prefix caching is on by default too, hashing token blocks so a repeated system prompt is prefilled once; the project's own documentation is careful to note that it only shortens prefill and does nothing for decode, which means it buys almost nothing on long generations with no shared prefix. Speculative decoding is available through n-gram, EAGLE and DFlash proposers rather than a separate small draft model.",{"type":123,"level":124,"id":97,"text":98},{"type":116,"content":167},[168],"Installation is one command, and the documented platform is Linux with Python 3.10 to 3.13. Serving one model looks like this:",{"type":170,"code":171},"code","uv venv --python 3.12 --seed\nsource .venv\u002Fbin\u002Factivate\nuv pip install vllm --torch-backend=auto\n\n# Prefix caching and chunked prefill are on by default in V1.\nvllm serve Qwen\u002FQwen3-8B \\\n  --served-model-name qwen3-8b \\\n  --max-model-len 32768 \\\n  --gpu-memory-utilization 0.9 \\\n  --tensor-parallel-size 2 \\\n  --api-key \"$VLLM_TOKEN\"",{"type":116,"content":173},[174,175,178,179,182,183,186,187,190,191,194,195,198],"The flags are mostly capacity decisions. ",{"tag":170,"children":176},[177],"--max-model-len"," caps context, and therefore how much KV cache a single request can hold, so set it to the longest prompt the product actually needs rather than the model maximum. ",{"tag":170,"children":180},[181],"--gpu-memory-utilization"," is the fraction of VRAM pre-allocated to weights and cache; anything left over after the weights is what the start-up profiling pass measures and turns into blocks. ",{"tag":170,"children":184},[185],"-O0"," through ",{"tag":170,"children":188},[189],"-O3"," control how hard the engine compiles and captures CUDA graphs, with ",{"tag":170,"children":192},[193],"-O2"," as the default; ",{"tag":170,"children":196},[197],"--enforce-eager"," skips both entirely. Talking to it needs no new client:",{"type":170,"code":200},"from openai import OpenAI\n\nclient = OpenAI(\n    api_key=\"EMPTY\",\n    base_url=\"http:\u002F\u002Flocalhost:8000\u002Fv1\",\n)\n\nstream = client.chat.completions.create(\n    model=\"qwen3-8b\",\n    messages=[\n        {\"role\": \"system\", \"content\": \"Answer in one sentence.\"},\n        {\"role\": \"user\", \"content\": \"Why beat a contiguous KV cache?\"},\n    ],\n    max_tokens=200,\n    temperature=0,\n    stream=True,\n)\n\nfor chunk in stream:\n    print(chunk.choices[0].delta.content or \"\", end=\"\", flush=True)",{"type":202,"variant":203,"title":204,"body":205},"callout","note","Sampling defaults come from the model, not from you",[206],[207,208,211,212,215],"By default the server applies the ",{"tag":170,"children":209},[210],"generation_config.json"," from the Hugging Face repository, so the model author's recommended sampling values quietly override what the request asks for. Pass ",{"tag":170,"children":213},[214],"--generation-config vllm"," to get the engine's own defaults, and pin sampling in the client when output has to stay stable across upgrades.",{"type":123,"level":124,"id":100,"text":101},{"type":116,"content":218},[219],"The famous numbers come from the 2023 launch post, not from recent benchmarks: LLaMA-7B on an A10G and LLaMA-13B on an A100 40GB, with request lengths sampled from ShareGPT, giving up to 24 times the throughput of Hugging Face Transformers and 2.2 to 3.5 times that of TGI. Those figures are old enough to be history rather than a specification, and the durable part of the story is the memory waste, not the multiples. What has changed since is that the floor moved: the serious competitors all adopted paged caches and in-flight batching, so the useful question is no longer who batches better but who tunes, patches and adds model support faster.",{"type":221,"head":222,"rows":229},"table",[223,225,227],[224],"Knob",[226],"What it changes",[228],"Move it when",[230,238,246,255,264,278],[231,234,236],[232],{"tag":170,"children":233},[177],[235],"Caps context, and with it the KV cache a single request can hold",[237],"The longest real prompt is far below the model maximum",[239,242,244],[240],{"tag":170,"children":241},[181],[243],"Fraction of VRAM pre-allocated to weights and KV cache",[245],"The start-up log reports a low block count, or requests start preempting",[247,251,253],[248],{"tag":170,"children":249},[250],"max_num_batched_tokens",[252],"Prefill tokens per step: small values favour inter-token latency, large values favour time to first token",[254],"Interactive chat versus offline batch work",[256,260,262],[257],{"tag":170,"children":258},[259],"--tensor-parallel-size",[261],"Splits weights across GPUs, which frees KV cache room on each",[263],"The model does not fit, or the KV cache is the binding constraint",[265,271,276],[266,268,269],{"tag":170,"children":267},[185]," to ",{"tag":170,"children":270},[189],[272,273,275],"Compilation and CUDA graph capture; ",{"tag":170,"children":274},[193]," is the default",[277],"Boot time matters more than steady-state decode",[279,282,284],[280],{"tag":170,"children":281},[197],[283],"Skips compilation and graph capture entirely",[285],"Development loops, or measuring how much of a boot is capture",{"type":116,"content":287},[288],"The failure mode that actually turns up in production is preemption. When the KV cache cannot hold every running sequence, vLLM preempts requests and recomputes them from the prompt once space returns, which barely registers in aggregate throughput and dominates tail latency. V1 defaults to recompute rather than swap precisely because swapping cost more, so the cure is capacity rather than a flag.",{"type":202,"variant":290,"title":291,"body":292},"warn","Recognise it in the log",[293],[294,295,298,299,302,303,306],"The scheduler warns with ",{"tag":170,"children":296},[297],"Sequence group 0 is preempted by PreemptionMode.RECOMPUTE mode because there is not enough KV cache space",". Set ",{"tag":170,"children":300},[301],"disable_log_stats=False"," to log the cumulative preemption count, or alert on ",{"tag":170,"children":304},[305],"vllm:kv_cache_usage_perc"," sitting near 1.0. Raising utilisation or tensor parallelism helps; pushing utilisation too far is how a co-tenant on the same host turns into an out-of-memory kill.",{"type":123,"level":124,"id":103,"text":104},{"type":123,"level":309,"id":310,"text":311},3,"observability","Observability comes first",{"type":116,"content":313},[314,315,318,319,322],"The engine exposes a Prometheus endpoint at ",{"tag":170,"children":316},[317],"\u002Fmetrics"," under a ",{"tag":170,"children":320},[321],"vllm:"," prefix. The metrics design document is unusually explicit about the intent: server-level gauges are there to explain the request-level histograms, and the request-level histograms are the series an operator is meant to alert on.",{"type":129,"ordered":130,"items":324},[325,330,335,340,348],[326,329],{"tag":170,"children":327},[328],"vllm:time_to_first_token_seconds"," — prefill cost, what a user feels on a cold prompt",[331,334],{"tag":170,"children":332},[333],"vllm:inter_token_latency_seconds"," — decode speed, what a user feels once generation has started",[336,339],{"tag":170,"children":337},[338],"vllm:e2e_request_latency_seconds"," — the series a timeout rule should be written against",[341,343,344,347],{"tag":170,"children":342},[305]," and ",{"tag":170,"children":345},[346],"vllm:num_requests_running"," — capacity; when both sit at their limits together, the queue is growing",[349,352,353,356],{"tag":170,"children":350},[351],"vllm:prefix_cache_queries"," against ",{"tag":170,"children":354},[355],"vllm:prefix_cache_hits"," — the ratio says whether shared prompts are actually reused, and therefore whether this workload suits the engine",{"type":123,"level":309,"id":358,"text":359},"process-and-cpu","Process count and CPU",{"type":116,"content":361},[362],"A four-GPU deployment is not one process. V1 runs one API server process, one engine core process and one worker process per GPU — six in total for a single node at tensor parallel size four — and a data-parallel deployment adds a coordinator on top. The tuning guide puts the floor at 2 plus N physical cores for N GPUs, because the engine core runs a busy loop and degrades visibly under CPU starvation.",{"type":202,"variant":364,"title":365,"body":366},"tip","Starved CPU looks like a slow GPU",[367],[368],"The guide names tokenisation, scheduling latency and streaming detokenisation as what suffers first, with lower-than-expected GPU utilisation as the symptom. With hyperthreading enabled, budget twice (2 + N) in vCPUs.",{"type":123,"level":309,"id":370,"text":371},"security","Security posture",{"type":116,"content":373},[374,375,378,379,382],"Authentication is one shared secret. The ",{"tag":170,"children":376},[377],"--api-key"," flag, or the ",{"tag":170,"children":380},[381],"VLLM_API_KEY"," environment variable, turns on a header check and accepts several keys at once so they can be rotated. There is no user model, no per-tenant quota and no authorisation layer, so anything that can reach the port can use the whole GPU.",{"type":202,"variant":290,"title":384,"body":385},"Development endpoints are a production incident waiting",[386],[387,388,391,392,395,396,343,399,402],"Setting ",{"tag":170,"children":389},[390],"VLLM_SERVER_DEV_MODE=1"," registers endpoints including ",{"tag":170,"children":393},[394],"\u002Fpause",", ",{"tag":170,"children":397},[398],"\u002Freset_prefix_cache",{"tag":170,"children":400},[401],"\u002Fcollective_rpc",", and the documentation carries its own security warning about them. Keep the port on a trusted network and put authentication, quotas and rate limiting in front of it.",{"type":123,"level":124,"id":106,"text":107},{"type":116,"content":405},[406,407,410,411,414],"Three things first. It is a GPU server, not a universal runtime: the documented target is Linux with CUDA or ROCm, so a CPU-only host or an Apple Silicon machine means a different project and a different model format. Boot time is real, because the default optimisation level compiles the model and captures CUDA graphs, and a cold container can spend minutes in the compiler before the first token. And the API is OpenAI-shaped rather than OpenAI-complete: the ",{"tag":170,"children":408},[409],"suffix"," parameter is unsupported, the ",{"tag":170,"children":412},[413],"user"," parameter is ignored, and parallel tool calls are best-effort and model-dependent.",{"type":221,"head":416,"rows":425},[417,419,421,423],[418],"Engine",[420],"Licence",[422],"Where it wins",[424],"What it costs you",[426,435,443,452],[427,430,431,433],[428],{"tag":134,"children":429},[23],[64],[432],"Breadth: the most architectures, the widest quantisation and hardware target set, one API surface for text, embeddings, rerank, speech and structured output",[434],"Compilation and graph capture on every boot, one model per server, and a busy loop that needs CPU behind it",[436,438,439,441],[437],"SGLang",[64],[440],"Radix-style prefix sharing and multi-turn state reuse, which suits agent and retrieval traffic that hits the same long context repeatedly",[442],"A smaller model zoo and a thinner serving surface outside chat completions",[444,446,448,450],[445],"TensorRT-LLM",[447],"Apache-2.0, NVIDIA stack only",[449],"Peak numbers on the newest NVIDIA parts, FP4 and FP8 kernels, first-class integration with Dynamo and Triton",[451],"NVIDIA only, and a rebuild whenever the model, the quantisation or the GPU generation changes",[453,455,457,459],[454],"llama.cpp",[456],"MIT",[458],"CPU, Apple Silicon and edge hardware, GGUF quantisation, a single binary with no Python runtime",[460],"A different model format, a weaker batched-serving story, no comparable surface for embeddings or rerank",{"type":116,"content":462},[463],"For most teams the real comparison is the first row against the second. vLLM and SGLang solve the same problem with the same primitives, both Apache-2.0, both OpenAI-compatible, and both will serve an ordinary chat workload well. SGLang's prefix tree is the better fit when the same long context is hit over and over; vLLM's model coverage and quantisation matrix are the better fit when new checkpoints arrive faster than workloads repeat.",{"type":123,"level":124,"id":109,"text":110},{"type":116,"content":466},[467],"vLLM is the tool to reach for by default, and the reason is not raw speed. It is surface area: the most models, the most quantisation formats, the most hardware targets, and one API that covers completions, chat, embeddings, rerank, speech and structured output. The costs are real and mostly boring to fix. Boot time is tunable, preemption is a capacity problem, and a starved CPU is a deployment mistake. None of them is a reason to pick something else.",{"type":129,"ordered":469,"items":470},true,[471,473,475,477,479,481],[472],"Pick it when you serve open-weight models on NVIDIA or AMD GPUs and the workload is batched: many concurrent requests rather than one at a time.",[474],"Pick it when model turnover is high, because a new checkpoint usually runs before the alternatives support it.",[476],"Pick it when the surrounding stack already speaks the OpenAI API, since the migration is a base URL.",[478],"Skip it for a laptop, an edge box or a single-user tool; llama.cpp or an MLX runtime does the job in a fraction of the memory.",[480],"Skip it when one model on one NVIDIA generation is committed and peak tokens per second is the only goal, because TensorRT-LLM will win that trade at the cost of the lock-in.",[482],"Benchmark SGLang before committing if the traffic is agentic or retrieval-heavy with heavy context reuse; the two are close enough that the deciding factor is usually which one the team can debug at 3am.",{"type":202,"variant":203,"title":484,"body":485},"The opinionated part",[486],[487],"The case for vLLM has quietly become a case about breadth rather than speed, and the projects that overtake it are unlikely to win on batching. They will win by being better at one workload. Teams are usually better served by one general engine plus one specialist than by an engine per workload, and that is the argument for standardising on vLLM even when a rival is measurably faster for a single traffic shape.",{"type":123,"level":124,"id":112,"text":113},{"type":129,"ordered":469,"items":490},[491,495,498,501,504,507,510,513,516],[492],{"tag":493,"href":28,"children":494},"a",[27],[496],{"tag":493,"href":31,"children":497},[30],[499],{"tag":493,"href":34,"children":500},[33],[502],{"tag":493,"href":37,"children":503},[36],[505],{"tag":493,"href":40,"children":506},[39],[508],{"tag":493,"href":43,"children":509},[42],[511],{"tag":493,"href":46,"children":512},[45],[514],{"tag":493,"href":49,"children":515},[48],[517],{"tag":493,"href":52,"children":518},[51],[520,567,650,712],{"slug":521,"published":522,"minutes":6,"category":7,"tags":523,"keywords":528,"about":535,"sources":539,"cover":558,"og":559,"expertise":55,"locales":560,"lang":57,"title":561,"description":562,"coverAlt":563,"url":564,"pricing":565,"kind":566},"ollama","2026-09-29",[524,525,454,526,527],"Local inference","Open models","GGUF","Model serving",[521,529,530,531,532,533,534],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[536],{"name":537,"url":538},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[540,543,546,549,552,555],{"title":541,"url":542},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":544,"url":545},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":547,"url":548},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":550,"url":551},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":553,"url":554},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":556,"url":557},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[57,58,59],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":568,"published":569,"minutes":6,"category":7,"tags":570,"keywords":576,"about":584,"sources":591,"cover":643,"og":644,"expertise":55,"locales":645,"lang":57,"title":646,"description":647,"coverAlt":648,"url":587,"pricing":649,"kind":571},"portkey","2026-09-28",[571,572,573,574,575],"LLM gateway","Guardrails","Routing","Observability","Cost control",[577,578,579,580,581,582,583],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[585,588],{"name":586,"url":587},"Portkey","https:\u002F\u002Fportkey.ai",{"name":589,"url":590},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[592,595,598,601,604,607,610,613,616,619,622,625,628,631,634,637,640],{"title":593,"url":594},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":596,"url":597},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":599,"url":600},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":602,"url":603},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":605,"url":606},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":608,"url":609},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":611,"url":612},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":614,"url":615},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":617,"url":618},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":620,"url":621},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":623,"url":624},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":626,"url":627},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":629,"url":630},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":632,"url":633},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":635,"url":636},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":638,"url":639},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":641,"url":642},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[57,58,59],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":651,"published":652,"minutes":653,"category":7,"tags":654,"keywords":660,"about":668,"sources":680,"cover":705,"og":706,"expertise":55,"locales":707,"lang":57,"title":708,"description":709,"coverAlt":710,"url":671,"pricing":711,"kind":655},"langfuse","2026-08-13",10,[655,656,657,658,659],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[651,661,662,663,664,665,666,667],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[669,672,674,677],{"name":670,"url":671},"Langfuse","https:\u002F\u002Flangfuse.com",{"name":657,"url":673},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":675,"url":676},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":678,"url":679},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[681,684,687,690,693,696,699,702],{"title":682,"url":683},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":685,"url":686},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":688,"url":689},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":691,"url":692},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":694,"url":695},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":697,"url":698},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":700,"url":701},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":703,"url":704},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[57,58,59],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":713,"published":714,"minutes":653,"category":7,"tags":715,"keywords":720,"about":727,"sources":731,"cover":747,"og":748,"expertise":55,"locales":749,"lang":57,"title":750,"description":751,"coverAlt":752,"url":753,"pricing":754,"kind":571},"openrouter","2026-07-23",[571,716,717,718,719],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[713,721,579,722,723,724,725,726],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[728],{"name":729,"url":730},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[732,735,736,739,742,745],{"title":733,"url":734},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":641,"url":642},{"title":737,"url":738},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":740,"url":741},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":743,"url":744},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":746,"url":730},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[57,58,59],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",1791383548974]