Tools/LLMOps & evals

Ollama review: the friendly way to run open models

Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.

Type
Local inference runtime
Pricing
MIT · free for personal use

··11 min read

  • Local inference
  • Open models
  • llama.cpp
  • GGUF
  • Model serving
Abstract cover art for the Ollama review

Key takeaways

  • Ollama’s own software is MIT licensed with no use restriction and no user threshold; the 700 million monthly active user limit people attribute to it belongs to Meta’s Llama Community Licence and applies to the weights, not the runtime.
  • Parallel requests share one context window rather than getting batched sequences, so memory scales as OLLAMA_NUM_PARALLEL times the context length and default concurrency is one.
  • The server on port 11434 has no authentication and the OpenAI-compatible endpoint ignores the key it asks for; it binds loopback by default and should stay there.
  • Ollama Cloud is a paid per-token service running alongside the runtime, with Pro at $20 a month and $60 of credits, while local inference on your own hardware stays free and unlimited.
  • It is the right tool for development, evaluation and single-node on-prem use, and the wrong one the moment throughput per GPU decides the project.

Ollama is a local inference runtime. It pulls open-weight models, places them on the GPU or CPU you have, and serves them over one HTTP API on port 11434. The vendor reports more than nine million installs a month, over a billion model downloads and 182,000 GitHub stars, and that reach is explicable: it is the least ceremony available for getting an open model to answer requests. The trade is stated in those same numbers. Ollama optimises for getting a model running, not for squeezing tokens out of a GPU, and a team that mistakes it for a production inference server will find that out.

In the stack it is a model server, not a framework. It replaces llama.cpp’s own server, LM Studio’s runtime and a hand-assembled Docker image, and it only competes with vLLM in the loosest possible sense. Above it sit LangChain, LlamaIndex, the coding agents and the self-hosted chat front ends; what makes Ollama interchangeable with all of them is the shape of its API, not anything it does inside it. What it does not offer is orchestration, continuous batching, autoscaling or multi-node tensor parallelism. Those still belong to vLLM or SGLang.

What it actually is

Ollama is a Go server with a model manager attached, not a model. It ships as one binary, one CLI and one Docker image, and the CLI is the whole product surface most people ever touch. The engines underneath are llama.cpp for CUDA, ROCm, Vulkan and CPU, and – since v0.40.0 – MLX as the default on Apple Silicon for the architectures it supports. Everything else is packaging around those engines, which is exactly why it is worth knowing where the boundary sits.

  • Licence: MIT for the software, with no use restriction, no user threshold and no additional clause.
  • Current release: v0.40.0, shipped as platform binaries and a Docker image. Cadence is fast but the numbers jump: v0.34.4 was followed by v0.35.1 and then straight to v0.40.0.
  • APIs: a native REST surface under /api, an OpenAI-compatible surface under /v1, and an Anthropic-compatible base URL, on the local server and on Ollama Cloud alike.
  • Engines: llama.cpp for CUDA, ROCm, Vulkan and CPU, plus MLX on Apple Silicon from v0.40.0. Nvidia needs compute capability 5.0 or newer; AMD needs the ROCm v7 driver on Linux.
  • Model format: GGUF for llama.cpp models, safetensors for MLX. Since v0.34.1, GGUF conversion has to be done with llama.cpp tooling rather than inside Ollama.
  • Customisation: a Modelfile with FROM, PARAMETER, TEMPLATE, SYSTEM, MESSAGE, LICENSE, REQUIRES and CAPABILITY, so a tuned model is a reviewable text file.
  • Also in the box: structured output against a JSON schema, tool calling, vision, embeddings, web search, experimental image generation on macOS, and decision models that return probabilities instead of text.

How it works

At the centre is an HTTP server that owns a model library and a VRAM-aware scheduler. A request names a model tag; if those weights are not already resident, the server loads them, and if they will not fit alongside what is already loaded, the request queues while an idle model is evicted. Loaded models stay resident for five minutes after the last request, context length is chosen from the VRAM tier the machine falls into, and parallel requests against one model share its context rather than each getting their own.

Request path through the Ollama serverA request arrives at /api/chat naming a model tag. The scheduler checks reported VRAM. If the weights are already resident, generation starts immediately; otherwise the request queues, weights are loaded into VRAM and the request is retried. After generation the model sits idle for the keep_alive window, five minutes by default, and is then unloaded so its VRAM is released.requestPOST /api/chatschedulerreads VRAMengineggml or MLXtokensstreamed NDJSONqueueMAX_QUEUEload weightsthen retryidlekeep_aliveunloadVRAM releasedno free VRAM5 min
The whole lifecycle of one request: a VRAM check, an immediate start if the weights are resident, a queue and load if not, then an idle window that ends in an unload.

That design has one consequence that catches people out: the model is the unit of scheduling and the unit of cost. Switching tags under load means a load, a VRAM check and, on a machine with one GPU, possibly the eviction of whatever was serving everyone else. Ollama’s answer is OLLAMA_MAX_LOADED_MODELS and OLLAMA_NUM_PARALLEL, and both of them are paid for in memory rather than in compute.

Getting started

Installation is a shell script, a desktop app or a Docker image. The API is one POST, and a local server needs no key and no configuration file. The snippet below is the smallest thing worth writing in production: a streaming call with an explicit context window, a model pinned resident between calls, and the timing fields the server already computes.

import json, urllib.request, time

URL = "http://localhost:11434/api/chat"
BODY = {
    "model": "gemma4",
    "messages": [{"role": "user", "content": "Summarise this ticket in one line."}],
    "stream": True,
    "keep_alive": "30m",              # keep the weights resident between calls
    "options": {"num_ctx": 8192, "temperature": 0},
}

request = urllib.request.Request(
    URL, data=json.dumps(BODY).encode(), headers={"Content-Type": "application/json"})

chunks = []
with urllib.request.urlopen(request) as response:
    for line in response:                       # newline-delimited JSON events
        event = json.loads(line)
        if "message" in event:
            chunks.append(event["message"].get("content", ""))
        if event.get("done"):
            seconds = event["eval_duration"] / 1e9   # nanoseconds
            print(f"{event['eval_count']} tokens in {seconds:.1f}s"
                  f" -> {event['eval_count'] / seconds:.1f} tok/s")
            print(f"prompt tokens {event['prompt_eval_count']}, cached"
                  f" {event.get('prompt_eval_cached_count', 0)},"
                  f" load {event['load_duration'] / 1e9:.1f}s")

print("".join(chunks))

Two fields are worth wiring into a dashboard: eval_count divided by eval_duration for generation speed, and load_duration for the cold penalty. prompt_eval_cached_count reports how many prompt tokens came from the cache, which is the only free speedup Ollama offers – put the stable prefix of a prompt first and the variable part last, and the shared system prompt stops being re-evaluated on every call.

The licence, and the clause that is not there

Ollama’s software is MIT licensed, full stop. The LICENSE file in the ollama/ollama repository is the unmodified MIT text: no additional clause, no user-count threshold, no use restriction. There is no OpenAI clause in it and no 700 million monthly active user limit, and that is worth saying plainly because the claim circulates widely. The 700 million threshold belongs to Meta’s Llama Community Licence, which covers Llama weights, not to the runtime that serves them. Ollama also raises money from paid cloud tiers and was funded with a $65 million round in July 2026, and neither fact puts a clause in the source licence.

The MIT grant covers the server binary. It says nothing about the models run through it, and that is where the real licence exposure sits. The library serves weights from a dozen publishers under a dozen terms, so the binding licence is the one on the model card, not the one on the runtime. A Modelfile records a LICENSE instruction alongside the weights, and ollama show --modelfile prints it – that string, per model version, is the artefact to keep in a compliance register.

  • MIT on the runtime. Use it commercially, fork it, ship it inside a product. No attribution duty beyond keeping the copyright notice, no revenue threshold, no user threshold, no telemetry obligation.
  • The model licence is separate. Meta’s Llama Community Licence requires a separate licence from Meta above 700 million monthly active users, and other families add their own revenue or user ceilings. Some models ship non-commercial terms. Read the model card.
  • ollama.com itself is not MIT. The hosted cloud, the library accounts and the paid tiers fall under the Terms of Service, last updated May 2026: binding arbitration in San Francisco, California law, a liability cap at amounts paid in the preceding twelve months, and a clause barring the use of the service to develop competing products.

Running it in production

Concurrency is where local runtimes are honest about their limits. Ollama runs one request per model by default, and parallel requests share a context rather than being batched independently: the documentation states it directly, a 2,000-token context with four parallel requests behaves as an 8,000-token context, and required RAM scales as parallel requests times context length. That is a workable trade for interactive use and a bad one for batch work.

SettingDefaultWhat it costs
OLLAMA_NUM_PARALLEL1RAM scales linearly; one shared context
OLLAMA_MAX_LOADED_MODELS3 per GPU, 3 on CPUFull VRAM for every resident model
OLLAMA_MAX_QUEUE512Queue depth before requests get a 503
OLLAMA_KEEP_ALIVE5 minutesIdle VRAM held; -1 pins it, 0 unloads now
OLLAMA_KV_CACHE_TYPEf16q8_0 halves it, q4_0 quarters it, with a precision cost

The server has no authentication. It binds 127.0.0.1 by default, and the OpenAI-compatible endpoint demands an API key value that it then ignores, which is a precise statement of how much the local surface is trusted. Anything that sets OLLAMA_HOST to a routable address is publishing an unauthenticated inference endpoint. In January 2026 researchers reported roughly 175,000 publicly reachable Ollama servers across 130 countries, most of them exposed by binding to 0.0.0.0.

  • Leave the bind address at 127.0.0.1. There is no auth to configure, so a reverse proxy with TLS and a real access check is the only access control on offer.
  • Set OLLAMA_ORIGINS explicitly if a browser client needs it. Loopback origins are already allowed by default, so the default is fine for local tools and wrong for anything shared.
  • On machines that must not reach ollama.com at all, set OLLAMA_NO_CLOUD=1 or disable_ollama_cloud in ~/.ollama/server.json, restart, and confirm the log line Ollama cloud disabled: true.
  • Remember that the desktop app registers as a login item on macOS and Windows, so port 11434 starts serving on boot whether anyone asked for it or not.

Where it shingiles

The weaknesses are real and they cluster in one place: throughput per GPU. One request per model by default, no independent batching of parallel requests, no multi-node parallelism and no tensor-parallel serving path. For an interactive endpoint with a handful of users that is invisible. For anything with a queue, a batch job or a cost target it is not: the same GPU returns a fraction of what vLLM returns on the same model, and the gap is not a configuration problem. It is the design.

OllamavLLMllama.cpp server
Installscript, app, imagepip, containerbinary or build
Throughput per GPUlow to mediumhighlow to medium
Independent batchingnoyesno
Hardware reachCUDA, ROCm, Vulkan, Metal, CPUCUDA, ROCmCUDA, Vulkan, Metal, CPU
Operational surfaceone server, env varsflags, metrics, clusterone binary, flags
LicenceMITApache 2.0MIT

Against LM Studio the comparison is close and is really about packaging: LM Studio has a GUI and a model browser, Ollama has a CLI, a Docker image and first-class headless use, and the Open WebUI ecosystem grew up around Ollama’s API shape. Against llama.cpp’s own server the difference is the model manager and the scheduler, which is worth a great deal to a team and nothing at all to somebody who already knows llama.cpp. Against vLLM the difference is the entire business case: if the question is how to serve this at a predictable cost, vLLM or SGLang is the tool and Ollama is the development environment to prototype in.

The other thing to weigh is direction. Ollama Cloud now carries Pro, Max and Team tiers, per-token model pricing and a model access control story aimed at companies, which means the project is a commercial inference provider as well as a runtime. That funds the release cadence and is not a criticism. But it means the centre of gravity is moving from running a model on your own machine to signing in to use a larger one, and a team that depends on local inference should own the version, the model pins and the artefacts rather than assume the surface stays where it is.

Verdict

Ollama is the best default answer to how to run an open model without hiring somebody to operate llama.cpp. It is MIT, it starts from one command, its API is the compatibility layer almost every agent framework already speaks, and it hides an enormous amount of GPU scheduling behind two environment variables. What it is not is a scalable inference platform, and the moment a request queue becomes visible on a dashboard, that is the signal to move rather than to tune.

  1. Use it for local development, evaluation harnesses, CI fixtures, on-prem installs where a handful of people share one workstation, and privacy-bound workloads where prompts must not leave the building.
  2. Use it for dropping a local or self-hosted endpoint into an agent framework, because the OpenAI-compatible surface removes the need for a provider-specific client.
  3. Use it for getting a team to a working open-model prototype in an afternoon. That is a genuine operational win and the reason most of its users never leave.
  4. Skip it for multi-user serving under load, batch generation, or anything where tokens per second per GPU is the metric the project is judged on.
  5. Skip it for regulated environments that need authentication, audit logging or a scheduler under the inference layer. Put a proxy in front, or use a different runtime.

Sources

  1. Ollama API documentation
  2. Ollama on GitHub, with the MIT LICENSE file
  3. Ollama terms of service, last updated May 2026
  4. Ollama pricing, cloud plans and per-token model rates
  5. Hardware support: Nvidia, AMD, Metal and Vulkan
  6. OpenAI compatibility, including what is not supported

Frequently asked questions

Is Ollama really MIT licensed, or is there a usage limit?

The software is MIT licensed with no additional clause, so there is no user-count threshold and no OpenAI restriction. The 700 million monthly active user clause people attribute to Ollama comes from Meta’s Llama Community Licence and applies to Llama weights you download, not to the runtime that serves them. Access to ollama.com itself is separately governed by the terms of service, last updated May 2026.

Does Ollama need an API key?

Not for a local server. The local instance has no authentication at all, and the OpenAI-compatible endpoint still demands a key value that it then ignores, so any placeholder works. Cloud requests to https://ollama.com do need a key, and a local server signed in with ollama signin can proxy to cloud models.

How do I serve more than one request at a time?

Set OLLAMA_NUM_PARALLEL, which defaults to 1. Memory scales with parallel requests times context length, and the context window is shared across them, so four parallel requests on a 2,000-token context behave as an 8,000-token context. Requests above the limit queue, up to OLLAMA_MAX_QUEUE, which defaults to 512 before a 503 is returned.

Is Ollama faster than vLLM?

No, and it is not trying to be. Ollama serves one request per model by default, batches parallel requests into a shared context rather than independently, and has no multi-node parallelism. For interactive use with a few concurrent callers the difference is small; for batch throughput per GPU it is the entire decision.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.