[{"data":1,"prerenderedAt":701},["ShallowReactive",2],{"tool-ollama-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":21,"sources":25,"cover":44,"og":45,"expertise":46,"locales":47,"lang":48,"title":51,"description":52,"coverAlt":53,"url":54,"pricing":55,"kind":56,"metaTitle":57,"takeaways":58,"faq":64,"toc":77,"blocks":102,"others":458},"ollama","2026-09-29",11,"llmops",[9,10,11,12,13],"Local inference","Open models","llama.cpp","GGUF","Model serving",[4,15,16,17,18,19,20],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[22],{"name":23,"url":24},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[26,29,32,35,38,41],{"title":27,"url":28},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":30,"url":31},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":33,"url":34},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":36,"url":37},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":39,"url":40},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":42,"url":43},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg","ai-engineer",[48,49,50],"en","de","hu","Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime","Ollama review: the friendly local runtime · Balázs Csorba",[59,60,61,62,63],"Ollama’s own software is MIT licensed with no use restriction and no user threshold; the 700 million monthly active user limit people attribute to it belongs to Meta’s Llama Community Licence and applies to the weights, not the runtime.","Parallel requests share one context window rather than getting batched sequences, so memory scales as OLLAMA_NUM_PARALLEL times the context length and default concurrency is one.","The server on port 11434 has no authentication and the OpenAI-compatible endpoint ignores the key it asks for; it binds loopback by default and should stay there.","Ollama Cloud is a paid per-token service running alongside the runtime, with Pro at $20 a month and $60 of credits, while local inference on your own hardware stays free and unlimited.","It is the right tool for development, evaluation and single-node on-prem use, and the wrong one the moment throughput per GPU decides the project.",[65,68,71,74],{"q":66,"a":67},"Is Ollama really MIT licensed, or is there a usage limit?","The software is MIT licensed with no additional clause, so there is no user-count threshold and no OpenAI restriction. The 700 million monthly active user clause people attribute to Ollama comes from Meta’s Llama Community Licence and applies to Llama weights you download, not to the runtime that serves them. Access to ollama.com itself is separately governed by the terms of service, last updated May 2026.",{"q":69,"a":70},"Does Ollama need an API key?","Not for a local server. The local instance has no authentication at all, and the OpenAI-compatible endpoint still demands a key value that it then ignores, so any placeholder works. Cloud requests to https:\u002F\u002Follama.com do need a key, and a local server signed in with ollama signin can proxy to cloud models.",{"q":72,"a":73},"How do I serve more than one request at a time?","Set OLLAMA_NUM_PARALLEL, which defaults to 1. Memory scales with parallel requests times context length, and the context window is shared across them, so four parallel requests on a 2,000-token context behave as an 8,000-token context. Requests above the limit queue, up to OLLAMA_MAX_QUEUE, which defaults to 512 before a 503 is returned.",{"q":75,"a":76},"Is Ollama faster than vLLM?","No, and it is not trying to be. Ollama serves one request per model by default, batches parallel requests into a shared context rather than independently, and has no multi-node parallelism. For interactive use with a few concurrent callers the difference is small; for batch throughput per GPU it is the entire decision.",[78,81,84,87,90,93,96,99],{"id":79,"title":80},"what-it-is","What it actually is",{"id":82,"title":83},"how-it-works","How it works",{"id":85,"title":86},"getting-started","Getting started",{"id":88,"title":89},"licence","The licence, and the clause that is not there",{"id":91,"title":92},"production","Running it in production",{"id":94,"title":95},"where-it-shingles","Where it shingiles",{"id":97,"title":98},"verdict","Verdict",{"id":100,"title":101},"sources","Sources",[103,107,110,113,116,156,157,160,169,172,173,176,179,182,211,212,215,218,235,240,241,244,313,316,326,331,332,335,397,400,403,404,407,432,436,437],{"type":104,"content":105},"paragraph",[106],"Ollama is a local inference runtime. It pulls open-weight models, places them on the GPU or CPU you have, and serves them over one HTTP API on port 11434. The vendor reports more than nine million installs a month, over a billion model downloads and 182,000 GitHub stars, and that reach is explicable: it is the least ceremony available for getting an open model to answer requests. The trade is stated in those same numbers. Ollama optimises for getting a model running, not for squeezing tokens out of a GPU, and a team that mistakes it for a production inference server will find that out.",{"type":104,"content":108},[109],"In the stack it is a model server, not a framework. It replaces llama.cpp’s own server, LM Studio’s runtime and a hand-assembled Docker image, and it only competes with vLLM in the loosest possible sense. Above it sit LangChain, LlamaIndex, the coding agents and the self-hosted chat front ends; what makes Ollama interchangeable with all of them is the shape of its API, not anything it does inside it. What it does not offer is orchestration, continuous batching, autoscaling or multi-node tensor parallelism. Those still belong to vLLM or SGLang.",{"type":111,"level":112,"id":79,"text":80},"heading",2,{"type":104,"content":114},[115],"Ollama is a Go server with a model manager attached, not a model. It ships as one binary, one CLI and one Docker image, and the CLI is the whole product surface most people ever touch. The engines underneath are llama.cpp for CUDA, ROCm, Vulkan and CPU, and – since v0.40.0 – MLX as the default on Apple Silicon for the architectures it supports. Everything else is packaging around those engines, which is exactly why it is worth knowing where the boundary sits.",{"type":117,"ordered":118,"items":119},"list",false,[120,126,131,136,141,146,151],[121,125],{"tag":122,"children":123},"strong",[124],"Licence:"," MIT for the software, with no use restriction, no user threshold and no additional clause.",[127,130],{"tag":122,"children":128},[129],"Current release:"," v0.40.0, shipped as platform binaries and a Docker image. Cadence is fast but the numbers jump: v0.34.4 was followed by v0.35.1 and then straight to v0.40.0.",[132,135],{"tag":122,"children":133},[134],"APIs:"," a native REST surface under \u002Fapi, an OpenAI-compatible surface under \u002Fv1, and an Anthropic-compatible base URL, on the local server and on Ollama Cloud alike.",[137,140],{"tag":122,"children":138},[139],"Engines:"," llama.cpp for CUDA, ROCm, Vulkan and CPU, plus MLX on Apple Silicon from v0.40.0. Nvidia needs compute capability 5.0 or newer; AMD needs the ROCm v7 driver on Linux.",[142,145],{"tag":122,"children":143},[144],"Model format:"," GGUF for llama.cpp models, safetensors for MLX. Since v0.34.1, GGUF conversion has to be done with llama.cpp tooling rather than inside Ollama.",[147,150],{"tag":122,"children":148},[149],"Customisation:"," a Modelfile with FROM, PARAMETER, TEMPLATE, SYSTEM, MESSAGE, LICENSE, REQUIRES and CAPABILITY, so a tuned model is a reviewable text file.",[152,155],{"tag":122,"children":153},[154],"Also in the box:"," structured output against a JSON schema, tool calling, vision, embeddings, web search, experimental image generation on macOS, and decision models that return probabilities instead of text.",{"type":111,"level":112,"id":82,"text":83},{"type":104,"content":158},[159],"At the centre is an HTTP server that owns a model library and a VRAM-aware scheduler. A request names a model tag; if those weights are not already resident, the server loads them, and if they will not fit alongside what is already loaded, the request queues while an idle model is evicted. Loaded models stay resident for five minutes after the last request, context length is chosen from the VRAM tier the machine falls into, and parallel requests against one model share its context rather than each getting their own.",{"type":161,"attrs":162,"inner":166,"caption":167},"diagram",{"viewBox":163,"role":164,"aria-labelledby":165},"0 0 720 250","img","ol-path-t ol-path-d","\u003Ctitle id=\"ol-path-t\">Request path through the Ollama server\u003C\u002Ftitle>\u003Cdesc id=\"ol-path-d\">A request arrives at \u002Fapi\u002Fchat naming a model tag. The scheduler checks reported VRAM. If the weights are already resident, generation starts immediately; otherwise the request queues, weights are loaded into VRAM and the request is retried. After generation the model sits idle for the keep_alive window, five minutes by default, and is then unloaded so its VRAM is released.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-ol\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Crect x=\"20\" y=\"50\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"95\" y=\"78\" text-anchor=\"middle\" class=\"d-text\">request\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"98\" text-anchor=\"middle\" class=\"d-small\">POST \u002Fapi\u002Fchat\u003C\u002Ftext>\u003Crect x=\"195\" y=\"50\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"270\" y=\"78\" text-anchor=\"middle\" class=\"d-text\">scheduler\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"98\" text-anchor=\"middle\" class=\"d-small\">reads VRAM\u003C\u002Ftext>\u003Crect x=\"370\" y=\"50\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"445\" y=\"78\" text-anchor=\"middle\" class=\"d-text\">engine\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"98\" text-anchor=\"middle\" class=\"d-small\">ggml or MLX\u003C\u002Ftext>\u003Crect x=\"545\" y=\"50\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"620\" y=\"78\" text-anchor=\"middle\" class=\"d-text\">tokens\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"98\" text-anchor=\"middle\" class=\"d-small\">streamed NDJSON\u003C\u002Ftext>\u003Crect x=\"20\" y=\"170\" width=\"150\" height=\"60\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"95\" y=\"196\" text-anchor=\"middle\" class=\"d-text\">queue\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"215\" text-anchor=\"middle\" class=\"d-small\">MAX_QUEUE\u003C\u002Ftext>\u003Crect x=\"195\" y=\"170\" width=\"150\" height=\"60\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"270\" y=\"196\" text-anchor=\"middle\" class=\"d-text\">load weights\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"215\" text-anchor=\"middle\" class=\"d-small\">then retry\u003C\u002Ftext>\u003Crect x=\"370\" y=\"170\" width=\"150\" height=\"60\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"445\" y=\"196\" text-anchor=\"middle\" class=\"d-text\">idle\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"215\" text-anchor=\"middle\" class=\"d-small\">keep_alive\u003C\u002Ftext>\u003Crect x=\"545\" y=\"170\" width=\"150\" height=\"60\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"620\" y=\"196\" text-anchor=\"middle\" class=\"d-text\">unload\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"215\" text-anchor=\"middle\" class=\"d-small\">VRAM released\u003C\u002Ftext>\u003Cpath d=\"M170 82H193\" class=\"d-line\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M345 82H368\" class=\"d-line-accent\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M520 82H543\" class=\"d-line\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M270 116V142H95V168\" class=\"d-line d-dash\" marker-end=\"url(#ah-ol)\" \u002F>\u003Ctext x=\"182\" y=\"138\" text-anchor=\"middle\" class=\"d-label\">no free VRAM\u003C\u002Ftext>\u003Cpath d=\"M170 200H193\" class=\"d-line\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M345 200H445V120\" class=\"d-line-accent\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M445 116V168\" class=\"d-line d-dash\" marker-end=\"url(#ah-ol)\" \u002F>\u003Ctext x=\"455\" y=\"146\" class=\"d-label\">5 min\u003C\u002Ftext>\u003Cpath d=\"M520 200H543\" class=\"d-line\" marker-end=\"url(#ah-ol)\" \u002F>\u003Cpath d=\"M620 116V168\" class=\"d-line d-dash\" marker-end=\"url(#ah-ol)\" \u002F>",[168],"The whole lifecycle of one request: a VRAM check, an immediate start if the weights are resident, a queue and load if not, then an idle window that ends in an unload.",{"type":104,"content":170},[171],"That design has one consequence that catches people out: the model is the unit of scheduling and the unit of cost. Switching tags under load means a load, a VRAM check and, on a machine with one GPU, possibly the eviction of whatever was serving everyone else. Ollama’s answer is OLLAMA_MAX_LOADED_MODELS and OLLAMA_NUM_PARALLEL, and both of them are paid for in memory rather than in compute.",{"type":111,"level":112,"id":85,"text":86},{"type":104,"content":174},[175],"Installation is a shell script, a desktop app or a Docker image. The API is one POST, and a local server needs no key and no configuration file. The snippet below is the smallest thing worth writing in production: a streaming call with an explicit context window, a model pinned resident between calls, and the timing fields the server already computes.",{"type":177,"code":178},"code","import json, urllib.request, time\n\nURL = \"http:\u002F\u002Flocalhost:11434\u002Fapi\u002Fchat\"\nBODY = {\n    \"model\": \"gemma4\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"Summarise this ticket in one line.\"}],\n    \"stream\": True,\n    \"keep_alive\": \"30m\",              # keep the weights resident between calls\n    \"options\": {\"num_ctx\": 8192, \"temperature\": 0},\n}\n\nrequest = urllib.request.Request(\n    URL, data=json.dumps(BODY).encode(), headers={\"Content-Type\": \"application\u002Fjson\"})\n\nchunks = []\nwith urllib.request.urlopen(request) as response:\n    for line in response:                       # newline-delimited JSON events\n        event = json.loads(line)\n        if \"message\" in event:\n            chunks.append(event[\"message\"].get(\"content\", \"\"))\n        if event.get(\"done\"):\n            seconds = event[\"eval_duration\"] \u002F 1e9   # nanoseconds\n            print(f\"{event['eval_count']} tokens in {seconds:.1f}s\"\n                  f\" -> {event['eval_count'] \u002F seconds:.1f} tok\u002Fs\")\n            print(f\"prompt tokens {event['prompt_eval_count']}, cached\"\n                  f\" {event.get('prompt_eval_cached_count', 0)},\"\n                  f\" load {event['load_duration'] \u002F 1e9:.1f}s\")\n\nprint(\"\".join(chunks))",{"type":104,"content":180},[181],"Two fields are worth wiring into a dashboard: eval_count divided by eval_duration for generation speed, and load_duration for the cold penalty. prompt_eval_cached_count reports how many prompt tokens came from the cache, which is the only free speedup Ollama offers – put the stable prefix of a prompt first and the variable part last, and the shared system prompt stops being re-evaluated on every call.",{"type":183,"variant":184,"body":185},"callout","tip",[186],[187,188,190,191,194,195,198,199,202,203,206,207,210],"If the calling code already speaks OpenAI, point base_url at http:\u002F\u002Flocalhost:11434\u002Fv1\u002F and leave the key as ",{"tag":177,"children":189},[4]," – the server ignores it. The compatible surface covers chat, streaming, JSON mode, tools and vision, but not logprobs, ",{"tag":177,"children":192},[193],"tool_choice",", ",{"tag":177,"children":196},[197],"logit_bias"," or image URLs. The native ",{"tag":177,"children":200},[201],"\u002Fapi\u002Fchat"," endpoint has carried ",{"tag":177,"children":204},[205],"logprobs"," and ",{"tag":177,"children":208},[209],"top_logprobs"," for longer, which is the reason to prefer it whenever token probabilities matter.",{"type":111,"level":112,"id":88,"text":89},{"type":104,"content":213},[214],"Ollama’s software is MIT licensed, full stop. The LICENSE file in the ollama\u002Follama repository is the unmodified MIT text: no additional clause, no user-count threshold, no use restriction. There is no OpenAI clause in it and no 700 million monthly active user limit, and that is worth saying plainly because the claim circulates widely. The 700 million threshold belongs to Meta’s Llama Community Licence, which covers Llama weights, not to the runtime that serves them. Ollama also raises money from paid cloud tiers and was funded with a $65 million round in July 2026, and neither fact puts a clause in the source licence.",{"type":104,"content":216},[217],"The MIT grant covers the server binary. It says nothing about the models run through it, and that is where the real licence exposure sits. The library serves weights from a dozen publishers under a dozen terms, so the binding licence is the one on the model card, not the one on the runtime. A Modelfile records a LICENSE instruction alongside the weights, and ollama show --modelfile prints it – that string, per model version, is the artefact to keep in a compliance register.",{"type":117,"ordered":118,"items":219},[220,225,230],[221,224],{"tag":122,"children":222},[223],"MIT on the runtime. ","Use it commercially, fork it, ship it inside a product. No attribution duty beyond keeping the copyright notice, no revenue threshold, no user threshold, no telemetry obligation.",[226,229],{"tag":122,"children":227},[228],"The model licence is separate. ","Meta’s Llama Community Licence requires a separate licence from Meta above 700 million monthly active users, and other families add their own revenue or user ceilings. Some models ship non-commercial terms. Read the model card.",[231,234],{"tag":122,"children":232},[233],"ollama.com itself is not MIT. ","The hosted cloud, the library accounts and the paid tiers fall under the Terms of Service, last updated May 2026: binding arbitration in San Francisco, California law, a liability cap at amounts paid in the preceding twelve months, and a clause barring the use of the service to develop competing products.",{"type":183,"variant":236,"body":237},"warn",[238],[239],"A pull is a download from someone else’s registry onto your disk, and that registry is not covered by the MIT grant. Teams with a model-approval process should mirror the GGUF files they actually use into a registry they control, because a tag is a mutable name and ollama pull follows it. Pin the version, record the licence, and keep a digest of what shipped.",{"type":111,"level":112,"id":91,"text":92},{"type":104,"content":242},[243],"Concurrency is where local runtimes are honest about their limits. Ollama runs one request per model by default, and parallel requests share a context rather than being batched independently: the documentation states it directly, a 2,000-token context with four parallel requests behaves as an 8,000-token context, and required RAM scales as parallel requests times context length. That is a workable trade for interactive use and a bad one for batch work.",{"type":245,"head":246,"rows":253},"table",[247,249,251],[248],"Setting",[250],"Default",[252],"What it costs",[254,265,274,285,302],[255,259,263],[256],{"tag":177,"children":257},[258],"OLLAMA_NUM_PARALLEL",[260],{"tag":177,"children":261},[262],"1",[264],"RAM scales linearly; one shared context",[266,270,272],[267],{"tag":177,"children":268},[269],"OLLAMA_MAX_LOADED_MODELS",[271],"3 per GPU, 3 on CPU",[273],"Full VRAM for every resident model",[275,279,283],[276],{"tag":177,"children":277},[278],"OLLAMA_MAX_QUEUE",[280],{"tag":177,"children":281},[282],"512",[284],"Queue depth before requests get a 503",[286,290,292],[287],{"tag":177,"children":288},[289],"OLLAMA_KEEP_ALIVE",[291],"5 minutes",[293,294,297,298,301],"Idle VRAM held; ",{"tag":177,"children":295},[296],"-1"," pins it, ",{"tag":177,"children":299},[300],"0"," unloads now",[303,307,311],[304],{"tag":177,"children":305},[306],"OLLAMA_KV_CACHE_TYPE",[308],{"tag":177,"children":309},[310],"f16",[312],"q8_0 halves it, q4_0 quarters it, with a precision cost",{"type":104,"content":314},[315],"The server has no authentication. It binds 127.0.0.1 by default, and the OpenAI-compatible endpoint demands an API key value that it then ignores, which is a precise statement of how much the local surface is trusted. Anything that sets OLLAMA_HOST to a routable address is publishing an unauthenticated inference endpoint. In January 2026 researchers reported roughly 175,000 publicly reachable Ollama servers across 130 countries, most of them exposed by binding to 0.0.0.0.",{"type":117,"ordered":118,"items":317},[318,320,322,324],[319],"Leave the bind address at 127.0.0.1. There is no auth to configure, so a reverse proxy with TLS and a real access check is the only access control on offer.",[321],"Set OLLAMA_ORIGINS explicitly if a browser client needs it. Loopback origins are already allowed by default, so the default is fine for local tools and wrong for anything shared.",[323],"On machines that must not reach ollama.com at all, set OLLAMA_NO_CLOUD=1 or disable_ollama_cloud in ~\u002F.ollama\u002Fserver.json, restart, and confirm the log line Ollama cloud disabled: true.",[325],"Remember that the desktop app registers as a login item on macOS and Windows, so port 11434 starts serving on boot whether anyone asked for it or not.",{"type":183,"variant":327,"body":328},"note",[329],[330],"Ollama Cloud is a different product from the local runtime, and the pricing page has moved well past free. Cloud models bill per million tokens out of usage credits – gemma4 lists $0.14 input and $0.40 output, gpt-oss:20b $0.07 and $0.30 – with off-peak rates outside 12:00 to 18:00 UTC on weekdays and all day at weekends. Pro is $20 a month with $60 of credits and three concurrent requests, Max is $100 with $300 and ten, Team is $500 in early access, and running models on your own hardware stays free and unlimited on every tier.",{"type":111,"level":112,"id":94,"text":95},{"type":104,"content":333},[334],"The weaknesses are real and they cluster in one place: throughput per GPU. One request per model by default, no independent batching of parallel requests, no multi-node parallelism and no tensor-parallel serving path. For an interactive endpoint with a handful of users that is invisible. For anything with a queue, a batch job or a cost target it is not: the same GPU returns a fraction of what vLLM returns on the same model, and the gap is not a configuration problem. It is the design.",{"type":245,"head":336,"rows":345},[337,339,341,343],[338],"",[340],"Ollama",[342],"vLLM",[344],"llama.cpp server",[346,355,363,371,380,389],[347,349,351,353],[348],"Install",[350],"script, app, image",[352],"pip, container",[354],"binary or build",[356,358,360,362],[357],"Throughput per GPU",[359],"low to medium",[361],"high",[359],[364,366,368,370],[365],"Independent batching",[367],"no",[369],"yes",[367],[372,374,376,378],[373],"Hardware reach",[375],"CUDA, ROCm, Vulkan, Metal, CPU",[377],"CUDA, ROCm",[379],"CUDA, Vulkan, Metal, CPU",[381,383,385,387],[382],"Operational surface",[384],"one server, env vars",[386],"flags, metrics, cluster",[388],"one binary, flags",[390,392,394,396],[391],"Licence",[393],"MIT",[395],"Apache 2.0",[393],{"type":104,"content":398},[399],"Against LM Studio the comparison is close and is really about packaging: LM Studio has a GUI and a model browser, Ollama has a CLI, a Docker image and first-class headless use, and the Open WebUI ecosystem grew up around Ollama’s API shape. Against llama.cpp’s own server the difference is the model manager and the scheduler, which is worth a great deal to a team and nothing at all to somebody who already knows llama.cpp. Against vLLM the difference is the entire business case: if the question is how to serve this at a predictable cost, vLLM or SGLang is the tool and Ollama is the development environment to prototype in.",{"type":104,"content":401},[402],"The other thing to weigh is direction. Ollama Cloud now carries Pro, Max and Team tiers, per-token model pricing and a model access control story aimed at companies, which means the project is a commercial inference provider as well as a runtime. That funds the release cadence and is not a criticism. But it means the centre of gravity is moving from running a model on your own machine to signing in to use a larger one, and a team that depends on local inference should own the version, the model pins and the artefacts rather than assume the surface stays where it is.",{"type":111,"level":112,"id":97,"text":98},{"type":104,"content":405},[406],"Ollama is the best default answer to how to run an open model without hiring somebody to operate llama.cpp. It is MIT, it starts from one command, its API is the compatibility layer almost every agent framework already speaks, and it hides an enormous amount of GPU scheduling behind two environment variables. What it is not is a scalable inference platform, and the moment a request queue becomes visible on a dashboard, that is the signal to move rather than to tune.",{"type":117,"ordered":408,"items":409},true,[410,415,419,423,428],[411,414],{"tag":122,"children":412},[413],"Use it for ","local development, evaluation harnesses, CI fixtures, on-prem installs where a handful of people share one workstation, and privacy-bound workloads where prompts must not leave the building.",[416,418],{"tag":122,"children":417},[413],"dropping a local or self-hosted endpoint into an agent framework, because the OpenAI-compatible surface removes the need for a provider-specific client.",[420,422],{"tag":122,"children":421},[413],"getting a team to a working open-model prototype in an afternoon. That is a genuine operational win and the reason most of its users never leave.",[424,427],{"tag":122,"children":425},[426],"Skip it for ","multi-user serving under load, batch generation, or anything where tokens per second per GPU is the metric the project is judged on.",[429,431],{"tag":122,"children":430},[426],"regulated environments that need authentication, audit logging or a scheduler under the inference layer. Put a proxy in front, or use a different runtime.",{"type":183,"variant":327,"body":433},[434],[435],"One closing caveat, stated plainly. “Runs locally” and “is MIT” are properties of two different artefacts. The runtime is MIT and auditable. The weights are whatever the publisher chose, the registry is a hosted service under its own terms, and a compliance review has to cover all three separately.",{"type":111,"level":112,"id":100,"text":101},{"type":117,"ordered":408,"items":438},[439,443,446,449,452,455],[440],{"tag":441,"href":28,"children":442},"a",[27],[444],{"tag":441,"href":31,"children":445},[30],[447],{"tag":441,"href":34,"children":448},[33],[450],{"tag":441,"href":37,"children":451},[36],[453],{"tag":441,"href":40,"children":454},[39],[456],{"tag":441,"href":43,"children":457},[42],[459,542,604,647],{"slug":460,"published":461,"minutes":6,"category":7,"tags":462,"keywords":468,"about":476,"sources":483,"cover":535,"og":536,"expertise":46,"locales":537,"lang":48,"title":538,"description":539,"coverAlt":540,"url":479,"pricing":541,"kind":463},"portkey","2026-09-28",[463,464,465,466,467],"LLM gateway","Guardrails","Routing","Observability","Cost control",[469,470,471,472,473,474,475],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[477,480],{"name":478,"url":479},"Portkey","https:\u002F\u002Fportkey.ai",{"name":481,"url":482},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[484,487,490,493,496,499,502,505,508,511,514,517,520,523,526,529,532],{"title":485,"url":486},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":488,"url":489},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":491,"url":492},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":494,"url":495},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":497,"url":498},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":500,"url":501},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":503,"url":504},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":506,"url":507},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":509,"url":510},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":512,"url":513},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":515,"url":516},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":518,"url":519},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":521,"url":522},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":524,"url":525},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":527,"url":528},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":530,"url":531},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":533,"url":534},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[48,49,50],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":543,"published":544,"minutes":545,"category":7,"tags":546,"keywords":552,"about":560,"sources":572,"cover":597,"og":598,"expertise":46,"locales":599,"lang":48,"title":600,"description":601,"coverAlt":602,"url":563,"pricing":603,"kind":547},"langfuse","2026-08-13",10,[547,548,549,550,551],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[543,553,554,555,556,557,558,559],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[561,564,566,569],{"name":562,"url":563},"Langfuse","https:\u002F\u002Flangfuse.com",{"name":549,"url":565},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":567,"url":568},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":570,"url":571},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[573,576,579,582,585,588,591,594],{"title":574,"url":575},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":577,"url":578},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":580,"url":581},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":583,"url":584},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":586,"url":587},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":589,"url":590},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":592,"url":593},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":595,"url":596},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[48,49,50],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":605,"published":606,"minutes":545,"category":7,"tags":607,"keywords":612,"about":619,"sources":623,"cover":639,"og":640,"expertise":46,"locales":641,"lang":48,"title":642,"description":643,"coverAlt":644,"url":645,"pricing":646,"kind":463},"openrouter","2026-07-23",[463,608,609,610,611],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[605,613,471,614,615,616,617,618],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[620],{"name":621,"url":622},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[624,627,628,631,634,637],{"title":625,"url":626},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":533,"url":534},{"title":629,"url":630},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":632,"url":633},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":635,"url":636},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":638,"url":622},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[48,49,50],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",{"slug":648,"published":649,"minutes":545,"category":7,"tags":650,"keywords":654,"about":661,"sources":669,"cover":693,"og":694,"expertise":46,"locales":695,"lang":48,"title":696,"description":697,"coverAlt":698,"url":664,"pricing":699,"kind":700},"braintrust","2026-07-07",[651,466,652,653,548],"Evaluations","LLM-as-a-judge","CI gates",[648,655,656,657,658,659,660],"braintrust pricing","braintrust eval","autoevals library","braintrust vs langfuse","llm evaluation platform","eval driven development",[662,665,668],{"name":663,"url":664},"Braintrust","https:\u002F\u002Fwww.braintrust.dev",{"name":666,"url":667},"Continuous integration","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FContinuous_integration",{"name":570,"url":571},[670,673,676,679,682,685,688,690],{"title":671,"url":672},"Braintrust pricing: plans, credits and usage rates","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing",{"title":674,"url":675},"Braintrust documentation: plans and limits","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fplans-and-limits",{"title":677,"url":678},"Braintrust documentation: evaluation quickstart","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs\u002Fevaluation-quickstart",{"title":680,"url":681},"Braintrust documentation: get started","https:\u002F\u002Fwww.braintrust.dev\u002Fdocs",{"title":683,"url":684},"GitHub: Braintrust organisation repositories","https:\u002F\u002Fgithub.com\u002Forgs\u002Fbraintrustdata\u002Frepositories",{"title":686,"url":687},"Arize Phoenix documentation: self-hosting","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix\u002Fself-hosting",{"title":689,"url":581},"Langfuse pricing: cloud plans and billable units",{"title":691,"url":692},"LangChain pricing: LangSmith plans","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fbraintrust\u002Fog.jpg",[48,49,50],"Braintrust review: eval-first observability with a hard meter","Braintrust turns production traces into datasets and gated experiments. What Starter and Pro really include, which parts are open source, and where Phoenix, Langfuse and LangSmith win.","A loop from instrumented application logs into a dataset, an experiment with scorers, and a comparison that gates the pull request before the change returns to the application.","Free · from $249 per month","Evaluation platform",1791383548672]