Tools/LLMOps & evals
vLLM reviewed for self-hosted inference
vLLM turns a Hugging Face checkpoint into an OpenAI-compatible server. What PagedAttention and continuous batching buy, and what running it actually costs.
- Type
- Inference server
- Pricing
- Apache-2.0
Balázs Csorba··11 min read
- Self-hosted inference
- OpenAI API
- PagedAttention
- GPU serving
- LLM runtime

Key takeaways
- vLLM is the default self-hosted serving engine for open-weight models on NVIDIA and AMD hardware, chosen for breadth of model, quantisation and API coverage rather than for being the fastest.
- PagedAttention and continuous batching are the reason it exists; the durable claim is that paged allocation cut KV cache waste from 60 to 80 per cent down to under 4 per cent, not the 24x throughput multiple from 2023.
- Prefix caching is enabled by default and only shortens prefill, so workloads with long repeated prompts gain the most and long generations with unique prompts gain nothing from it.
- The production failure mode is preemption by recompute when the KV cache is undersized; watch KV cache usage and the cumulative preemption count rather than aggregate throughput.
- A four-GPU deployment runs six processes and needs at least six physical CPU cores, which is the most common reason throughput lands below expectations.
vLLM is an open-source inference engine that turns a Hugging Face checkpoint into an HTTP server that speaks the OpenAI API. It is the layer between a model and everything that calls it, and its entire job is to keep the GPU busy. The verdict up front: on NVIDIA or AMD hardware running open-weight models it is the default choice, because no other project covers this much of the surface with this little glue code. On a laptop, on a CPU-only host, or with one model and one user, it is far too much machinery.
It competes in the serving layer, not the model layer. The alternatives are SGLang, which shares most of its design; NVIDIA's TensorRT-LLM, which trades breadth for peak numbers on NVIDIA parts; llama.cpp, which starts where vLLM gives up; and Hugging Face TGI, whose last release was v3.3.7 in December 2025. It is also the usual substitute for a hosted API, because the same OpenAI client code that talks to a provider can be pointed at a local process.
What it actually is
The facts that matter when picking a serving engine, all of them taken from the project's own documentation rather than from a vendor page:
- Apache-2.0, with no paid tier and no control plane hosted by the vendor.
- Started at UC Berkeley's Sky Computing Lab; the documentation credits more than 2,000 contributors and calls it one of the most active open-source AI projects.
- More than 200 model architectures on Hugging Face, spanning decoder-only, mixture-of-experts, hybrid attention, multimodal, embedding, rerank and reward models.
- An OpenAI-compatible server plus Anthropic Messages and Cohere embed and rerank endpoints, and it also covers speech and structured output. One model per server process.
- CUDA and ROCm as first-class targets, Intel XPU and Google TPU supported, with hardware plugins for Ascend NPUs, Gaudi, Spyre, Apple Silicon, MetaX and others.
- Quantisation across FP8, NVFP4, MXFP4, INT8 and INT4, plus GPTQ, AWQ, GGUF and compressed-tensors checkpoints.
- A minor release roughly every two weeks: v0.24.0 shipped on 29 June 2026, six weeks after v0.20.2 on 10 May.
How it works
Two ideas carry most of the throughput. PagedAttention stores the key and value cache in fixed-size blocks and maps them onto non-contiguous GPU memory through a block table, the way an operating system maps pages; the 2023 paper measured 60 to 80 per cent of KV cache wasted through fragmentation and over-reservation in earlier systems, against under 4 per cent for paged allocation. Continuous batching then admits new requests at every decode step instead of waiting for a batch to fill, which is where most of the latency improvement comes from.
The V1 engine, which replaced V0 during 2025, mixes prefill and decode in the same step and gives decode priority, so a long prompt no longer stalls streaming traffic. Chunked prefill is on by default. Prefix caching is on by default too, hashing token blocks so a repeated system prompt is prefilled once; the project's own documentation is careful to note that it only shortens prefill and does nothing for decode, which means it buys almost nothing on long generations with no shared prefix. Speculative decoding is available through n-gram, EAGLE and DFlash proposers rather than a separate small draft model.
Getting a server up
Installation is one command, and the documented platform is Linux with Python 3.10 to 3.13. Serving one model looks like this:
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
# Prefix caching and chunked prefill are on by default in V1.
vllm serve Qwen/Qwen3-8B \
--served-model-name qwen3-8b \
--max-model-len 32768 \
--gpu-memory-utilization 0.9 \
--tensor-parallel-size 2 \
--api-key "$VLLM_TOKEN"The flags are mostly capacity decisions. --max-model-len caps context, and therefore how much KV cache a single request can hold, so set it to the longest prompt the product actually needs rather than the model maximum. --gpu-memory-utilization is the fraction of VRAM pre-allocated to weights and cache; anything left over after the weights is what the start-up profiling pass measures and turns into blocks. -O0 through -O3 control how hard the engine compiles and captures CUDA graphs, with -O2 as the default; --enforce-eager skips both entirely. Talking to it needs no new client:
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:8000/v1",
)
stream = client.chat.completions.create(
model="qwen3-8b",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "Why beat a contiguous KV cache?"},
],
max_tokens=200,
temperature=0,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)What the throughput claims are worth
The famous numbers come from the 2023 launch post, not from recent benchmarks: LLaMA-7B on an A10G and LLaMA-13B on an A100 40GB, with request lengths sampled from ShareGPT, giving up to 24 times the throughput of Hugging Face Transformers and 2.2 to 3.5 times that of TGI. Those figures are old enough to be history rather than a specification, and the durable part of the story is the memory waste, not the multiples. What has changed since is that the floor moved: the serious competitors all adopted paged caches and in-flight batching, so the useful question is no longer who batches better but who tunes, patches and adds model support faster.
| Knob | What it changes | Move it when |
|---|---|---|
--max-model-len | Caps context, and with it the KV cache a single request can hold | The longest real prompt is far below the model maximum |
--gpu-memory-utilization | Fraction of VRAM pre-allocated to weights and KV cache | The start-up log reports a low block count, or requests start preempting |
max_num_batched_tokens | Prefill tokens per step: small values favour inter-token latency, large values favour time to first token | Interactive chat versus offline batch work |
--tensor-parallel-size | Splits weights across GPUs, which frees KV cache room on each | The model does not fit, or the KV cache is the binding constraint |
-O0 to -O3 | Compilation and CUDA graph capture; -O2 is the default | Boot time matters more than steady-state decode |
--enforce-eager | Skips compilation and graph capture entirely | Development loops, or measuring how much of a boot is capture |
The failure mode that actually turns up in production is preemption. When the KV cache cannot hold every running sequence, vLLM preempts requests and recomputes them from the prompt once space returns, which barely registers in aggregate throughput and dominates tail latency. V1 defaults to recompute rather than swap precisely because swapping cost more, so the cure is capacity rather than a flag.
Running it in production
Observability comes first
The engine exposes a Prometheus endpoint at /metrics under a vllm: prefix. The metrics design document is unusually explicit about the intent: server-level gauges are there to explain the request-level histograms, and the request-level histograms are the series an operator is meant to alert on.
vllm:time_to_first_token_seconds— prefill cost, what a user feels on a cold promptvllm:inter_token_latency_seconds— decode speed, what a user feels once generation has startedvllm:e2e_request_latency_seconds— the series a timeout rule should be written againstvllm:kv_cache_usage_percandvllm:num_requests_running— capacity; when both sit at their limits together, the queue is growingvllm:prefix_cache_queriesagainstvllm:prefix_cache_hits— the ratio says whether shared prompts are actually reused, and therefore whether this workload suits the engine
Process count and CPU
A four-GPU deployment is not one process. V1 runs one API server process, one engine core process and one worker process per GPU — six in total for a single node at tensor parallel size four — and a data-parallel deployment adds a coordinator on top. The tuning guide puts the floor at 2 plus N physical cores for N GPUs, because the engine core runs a busy loop and degrades visibly under CPU starvation.
Security posture
Authentication is one shared secret. The --api-key flag, or the VLLM_API_KEY environment variable, turns on a header check and accepts several keys at once so they can be rotated. There is no user model, no per-tenant quota and no authorisation layer, so anything that can reach the port can use the whole GPU.
Where it stings
Three things first. It is a GPU server, not a universal runtime: the documented target is Linux with CUDA or ROCm, so a CPU-only host or an Apple Silicon machine means a different project and a different model format. Boot time is real, because the default optimisation level compiles the model and captures CUDA graphs, and a cold container can spend minutes in the compiler before the first token. And the API is OpenAI-shaped rather than OpenAI-complete: the suffix parameter is unsupported, the user parameter is ignored, and parallel tool calls are best-effort and model-dependent.
| Engine | Licence | Where it wins | What it costs you |
|---|---|---|---|
| vLLM | Apache-2.0 | Breadth: the most architectures, the widest quantisation and hardware target set, one API surface for text, embeddings, rerank, speech and structured output | Compilation and graph capture on every boot, one model per server, and a busy loop that needs CPU behind it |
| SGLang | Apache-2.0 | Radix-style prefix sharing and multi-turn state reuse, which suits agent and retrieval traffic that hits the same long context repeatedly | A smaller model zoo and a thinner serving surface outside chat completions |
| TensorRT-LLM | Apache-2.0, NVIDIA stack only | Peak numbers on the newest NVIDIA parts, FP4 and FP8 kernels, first-class integration with Dynamo and Triton | NVIDIA only, and a rebuild whenever the model, the quantisation or the GPU generation changes |
| llama.cpp | MIT | CPU, Apple Silicon and edge hardware, GGUF quantisation, a single binary with no Python runtime | A different model format, a weaker batched-serving story, no comparable surface for embeddings or rerank |
For most teams the real comparison is the first row against the second. vLLM and SGLang solve the same problem with the same primitives, both Apache-2.0, both OpenAI-compatible, and both will serve an ordinary chat workload well. SGLang's prefix tree is the better fit when the same long context is hit over and over; vLLM's model coverage and quantisation matrix are the better fit when new checkpoints arrive faster than workloads repeat.
Verdict
vLLM is the tool to reach for by default, and the reason is not raw speed. It is surface area: the most models, the most quantisation formats, the most hardware targets, and one API that covers completions, chat, embeddings, rerank, speech and structured output. The costs are real and mostly boring to fix. Boot time is tunable, preemption is a capacity problem, and a starved CPU is a deployment mistake. None of them is a reason to pick something else.
- Pick it when you serve open-weight models on NVIDIA or AMD GPUs and the workload is batched: many concurrent requests rather than one at a time.
- Pick it when model turnover is high, because a new checkpoint usually runs before the alternatives support it.
- Pick it when the surrounding stack already speaks the OpenAI API, since the migration is a base URL.
- Skip it for a laptop, an edge box or a single-user tool; llama.cpp or an MLX runtime does the job in a fraction of the memory.
- Skip it when one model on one NVIDIA generation is committed and peak tokens per second is the only goal, because TensorRT-LLM will win that trade at the cost of the lock-in.
- Benchmark SGLang before committing if the traffic is agentic or retrieval-heavy with heavy context reuse; the two are close enough that the deciding factor is usually which one the team can debug at 3am.
Sources
- vLLM documentation
- Optimization and tuning for the V1 engine
- Architecture overview and the V1 process layout
- Metrics design
- Automatic prefix caching and its documented limits
- Online serving and the HTTP API surface
- Inside vLLM: anatomy of a high-throughput inference system
- vLLM: easy, fast and cheap LLM serving with PagedAttention
- Release history
Frequently asked questions
Is vLLM faster than SGLang or TensorRT-LLM?
Not categorically, and most published comparisons are not comparable. The 2023 vLLM launch post reported up to 24 times the throughput of Hugging Face Transformers and 2.2 to 3.5 times that of TGI, measured on LLaMA-7B and LLaMA-13B with request lengths sampled from ShareGPT. That predates competitors adopting paged caches, so benchmark your own request-length distribution rather than trusting any single multiple.
Do I need a GPU to run vLLM?
Linux with CUDA or ROCm is the documented target, with Python 3.10 to 3.13. Intel XPU and Google TPU are supported separately, and Apple Silicon is covered by a different, MLX-based project rather than by vLLM itself. CPU-only inference is llama.cpp territory, not vLLM's.
How much VRAM does vLLM need?
vLLM takes a configurable fraction of VRAM for weights and KV cache, then measures what is left by running a profiling forward pass at start-up and reports how many KV cache blocks fit. In practice the binding constraint is the model weights at your chosen quantisation; the remainder is KV cache, and how much that is decides your concurrency.
Can one vLLM server host several models?
Not natively. A server process hosts one model at a time, so multi-model deployments run one process per model behind a router, or use data parallelism to replicate the same model across several GPUs. LoRA adapters are the exception: they can be loaded and unloaded at runtime, and the documentation restricts that to local development.
How do I keep vLLM output reproducible across upgrades?
Pin the engine version, since minor releases land roughly every two weeks, and pass --generation-config vllm to stop the server applying the model's generation_config.json from Hugging Face, which otherwise overrides your sampling defaults silently. For deterministic runs, set temperature to zero and pin the model revision as well.