[{"data":1,"prerenderedAt":863},["ShallowReactive",2],{"tool-llama-cpp-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":34,"cover":76,"og":77,"expertise":78,"locales":79,"lang":80,"title":83,"description":84,"coverAlt":85,"url":24,"pricing":86,"kind":87,"metaTitle":88,"takeaways":89,"faq":95,"toc":108,"blocks":135,"others":508},"llama-cpp","2026-10-05",8,"llmops",[9,10,11,12,13],"llama.cpp","GGUF","Quantisation","Local inference","llama-server",[9,15,16,17,18,19,20,21],"llama.cpp review","GGUF quantisation levels","llama-server OpenAI compatible","llama.cpp vs Ollama","run LLM locally GDPR","llama.cpp CUDA Metal Vulkan","Q4_K_M vs Q8_0",[23,25,28,31],{"name":9,"url":24},"https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp",{"name":26,"url":27},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",{"name":29,"url":30},"Quantization (signal processing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FQuantization_(signal_processing)",{"name":32,"url":33},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",[35,37,40,43,46,49,52,55,58,61,64,67,70,73],{"title":36,"url":24},"llama.cpp repository: goals, backends, licence",{"title":38,"url":39},"llama.cpp releases: builds b11538 to b11541","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Freleases",{"title":41,"url":42},"llama.cpp build documentation: CMake flags","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Fdocs\u002Fbuild.md",{"title":44,"url":45},"llama.cpp server README: endpoints, slots and grammars","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fserver\u002FREADME.md",{"title":47,"url":48},"llama.cpp quantisation README: bits, size and speed","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fquantize\u002FREADME.md",{"title":50,"url":51},"GGUF specification in the ggml repository","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fggml\u002Fblob\u002Fmaster\u002Fdocs\u002Fgguf.md",{"title":53,"url":54},"Performance of llama.cpp on Apple Silicon M-series","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F4167",{"title":56,"url":57},"Performance of llama.cpp on Nvidia CUDA","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F15013",{"title":59,"url":60},"Ollama README: supported backends and REST API","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":62,"url":63},"LM Studio documentation: app overview","https:\u002F\u002Flmstudio.ai\u002Fdocs\u002Fapp",{"title":65,"url":66},"vLLM README: features, hardware and licence","https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm",{"title":68,"url":69},"vLLM documentation: GGUF support","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Flatest\u002Ffeatures\u002Fquantization\u002Fgguf.html",{"title":71,"url":72},"GDPR Article 28: processor","https:\u002F\u002Fgdpr-info.eu\u002Fart-28-gdpr\u002F",{"title":74,"url":75},"GDPR Article 32: security of processing","https:\u002F\u002Fgdpr-info.eu\u002Fart-32-gdpr\u002F","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fog.jpg","ai-engineer",[80,81,82],"en","de","hu","llama.cpp review: the local engine under Ollama and LM Studio","llama.cpp runs open models in plain C and C++ on Metal, CUDA, Vulkan or the CPU. I cover GGUF quants, llama-server and where it falls short.","Cover art for the llama.cpp review: one GGUF file fans out to Metal, CUDA, Vulkan and CPU, then one OpenAI-style API.","MIT · free","Local inference runtime","llama.cpp review: GGUF, quants and the server · Balázs Csorba",[90,91,92,93,94],"llama.cpp is the MIT-licensed engine under Ollama and LM Studio. Use it directly when you want the model file, the quant and every server flag under your control.","A GGUF file carries the weights, the tokeniser and the metadata. The build flag picks the backend, so one file runs on Metal, CUDA, Vulkan or the CPU.","Q4_K_M is the sensible starting quant, but the quantisation table measures size and speed only, so the quality check is yours to run.","Speed on Apple chips follows memory bandwidth, though not in a straight line. A high-end chip generates tens of tokens a second on a 7B model, and a data-centre GPU is roughly three times faster in the same kind of test.","Local inference keeps prompts on hardware you control, which takes a processor out of the inference step, but it does not remove your GDPR duties for logs, access and retention.",[96,99,102,105],{"q":97,"a":98},"Is llama.cpp free for commercial use?","The project is MIT-licensed, so you can use and ship it without a licence fee. The model weights you load keep their own licences, which can restrict commercial use, so check each model card before you ship a product on top of it.",{"q":100,"a":101},"What is the difference between llama.cpp and Ollama?","Ollama lists llama.cpp among its supported backends and adds a REST API for running and managing models. llama.cpp is the engine itself: you choose the GGUF file, the quant and every flag, and you update the binary yourself.",{"q":103,"a":104},"Which GGUF quant should I start with?","Q4_K_M. The quantisation README uses a naive Q4_K_M quantisation as its example. Move up to Q5_K_M or Q6_K if your own evals show a quality gap, because the speed table says nothing about quality.",{"q":106,"a":107},"Does a local llama.cpp server help with GDPR?","It helps because no model vendor receives the prompts, so that processor drops out of the chain. Access control, log retention and a lawful basis still apply. A rented GPU host that handles personal data for you is a processor and needs an Article 28 contract.",[109,112,115,118,121,123,126,129,132],{"id":110,"title":111},"what-it-is","What it is",{"id":113,"title":114},"how-it-works","How it works",{"id":116,"title":117},"getting-started","Getting started",{"id":119,"title":120},"quantisation-and-speed","Quantisation and speed",{"id":13,"title":122},"llama-server: the API, slots and structured output",{"id":124,"title":125},"cost-and-deployment","Cost and deployment",{"id":127,"title":128},"where-it-falls-short","Where it falls short",{"id":130,"title":131},"verdict","Verdict",{"id":133,"title":134},"sources","Sources",[136,140,143,146,174,175,178,187,188,191,229,232,235,236,239,304,313,320,323,326,332,333,336,339,342,344,345,381,384,392,398,399,426,427,430,462,463],{"type":137,"content":138},"paragraph",[139],"llama.cpp is the C and C++ engine that runs open-weight language models on hardware you control, and several friendlier local tools sit on top of it, Ollama and LM Studio among them. The verdict up front: use it directly when you want to choose the model file, the quantisation and every server flag yourself, on a laptop or a server you run. Skip it if you want one command that downloads and manages models for you, and skip it as the serving layer for a busy multi-user GPU service, where vLLM is the better fit.",{"type":141,"level":142,"id":110,"text":111},"heading",2,{"type":137,"content":144},[145],"The project states its goal as LLM and VLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. It is a plain C and C++ implementation with no external dependencies, built on the ggml tensor library. The newest build on the releases page is b11541, published on 10 October 2026.",{"type":147,"ordered":148,"items":149},"list",false,[150,156,161,165,169],[151,155],{"tag":152,"children":153},"strong",[154],"MIT licence"," for the whole project, so you can use and ship it without a licence fee.",[157,160],{"tag":152,"children":158},[159],"Backends"," for Apple Metal, NVIDIA CUDA, AMD HIP, Vulkan, OpenCL, SYCL and WebGPU, plus x86 and ARM CPU code paths.",[162,164],{"tag":152,"children":163},[13],", an OpenAI-compatible HTTP server with a built-in web UI.",[166,168],{"tag":152,"children":167},[11]," from 1.5 to 8 bits, with GGUF as the model format.",[170,173],{"tag":152,"children":171},[172],"Hugging Face support"," through the -hf flag, plus conversion scripts such as convert_hf_to_gguf.py.",{"type":141,"level":142,"id":113,"text":114},{"type":137,"content":176},[177],"The GGUF file does most of the work. The specification calls GGUF “a file format for storing models for inference with GGML and executors based on GGML”, designed for fast loading and saving. Each file holds a header with the tensor and metadata counts, typed key-value metadata, a description of every tensor (name, shape, type and offset) and the tensor data, padded to an alignment boundary. The spec lists memory-map compatibility as a goal, so the operating system can map the weights rather than copy them. The tokeniser travels in the file as well, under tokenizer.ggml keys. The spec warns that the embedded vocabulary may be less accurate than the original tokeniser, so check output quality after you convert a model.",{"type":179,"attrs":180,"inner":184,"caption":185},"diagram",{"viewBox":181,"role":182,"aria-labelledby":183},"0 0 720 330","img","d1-lc-t d1-lc-d","\u003Ctitle id=\"d1-lc-t\">One GGUF file, one API\u003C\u002Ftitle>\u003Cdesc id=\"d1-lc-d\">A GGUF file holds the weights, the tokeniser and the metadata. llama.cpp loads it and runs it on the backend chosen at build time, such as Metal, CUDA, Vulkan or the CPU. llama-server serves the model through an OpenAI-compatible API, which any OpenAI client can call. Only the backend changes with the hardware; the file and the API stay the same.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-lc\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">One GGUF file, one API\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">llama.cpp on your machine\u003C\u002Ftext>\u003Crect x=\"20\" y=\"60\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"95\" y=\"88\" text-anchor=\"middle\" class=\"d-text\">GGUF file\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"110\" text-anchor=\"middle\" class=\"d-small\">weights, tokeniser\u003C\u002Ftext>\u003Cpath d=\"M170 91 H193\" class=\"d-line\" marker-end=\"url(#ah-lc)\" \u002F>\u003Crect x=\"195\" y=\"60\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"270\" y=\"88\" text-anchor=\"middle\" class=\"d-text\">llama.cpp\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"110\" text-anchor=\"middle\" class=\"d-small\">ggml graph\u003C\u002Ftext>\u003Cpath d=\"M345 91 H368\" class=\"d-line\" marker-end=\"url(#ah-lc)\" \u002F>\u003Crect x=\"370\" y=\"60\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"445\" y=\"88\" text-anchor=\"middle\" class=\"d-text\">llama-server\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"110\" text-anchor=\"middle\" class=\"d-small\">OpenAI-style API\u003C\u002Ftext>\u003Cpath d=\"M520 91 H543\" class=\"d-line\" marker-end=\"url(#ah-lc)\" \u002F>\u003Crect x=\"545\" y=\"60\" width=\"150\" height=\"62\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"620\" y=\"88\" text-anchor=\"middle\" class=\"d-text\">Your app\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"110\" text-anchor=\"middle\" class=\"d-small\">any OpenAI client\u003C\u002Ftext>\u003Cpath d=\"M270 122 V180\" class=\"d-line\" \u002F>\u003Cpath d=\"M95 180 H620\" class=\"d-line\" \u002F>\u003Cpath d=\"M95 180 V200\" class=\"d-line\" \u002F>\u003Cpath d=\"M270 180 V200\" class=\"d-line\" \u002F>\u003Cpath d=\"M445 180 V200\" class=\"d-line\" \u002F>\u003Cpath d=\"M620 180 V200\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"200\" width=\"150\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"95\" y=\"224\" text-anchor=\"middle\" class=\"d-text\">Metal\u003C\u002Ftext>\u003Ctext x=\"95\" y=\"244\" text-anchor=\"middle\" class=\"d-small\">Apple Silicon\u003C\u002Ftext>\u003Crect x=\"195\" y=\"200\" width=\"150\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"270\" y=\"224\" text-anchor=\"middle\" class=\"d-text\">CUDA\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"244\" text-anchor=\"middle\" class=\"d-small\">NVIDIA GPUs\u003C\u002Ftext>\u003Crect x=\"370\" y=\"200\" width=\"150\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"445\" y=\"224\" text-anchor=\"middle\" class=\"d-text\">Vulkan\u003C\u002Ftext>\u003Ctext x=\"445\" y=\"244\" text-anchor=\"middle\" class=\"d-small\">GPU drivers\u003C\u002Ftext>\u003Crect x=\"545\" y=\"200\" width=\"150\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"620\" y=\"224\" text-anchor=\"middle\" class=\"d-text\">CPU\u003C\u002Ftext>\u003Ctext x=\"620\" y=\"244\" text-anchor=\"middle\" class=\"d-small\">x86 and ARM\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"290\" class=\"d-small\">Build flags pick the backend: -DGGML_CUDA=ON, -DGGML_VULKAN=ON, Metal by default on macOS.\u003C\u002Ftext>",[186],"One GGUF file and one API: the backend follows your hardware, and the interface stays the same.",{"type":141,"level":142,"id":116,"text":117},{"type":137,"content":189},[190],"The build documentation enables Metal by default on macOS, so a plain build on a Mac already includes the Apple backend. The other backends need a flag when you configure the build.",{"type":192,"head":193,"rows":200},"table",[194,196,198],[195],"Backend",[197],"Hardware",[199],"How to enable",[201,208,215,222],[202,204,206],[203],"Metal",[205],"Apple Silicon",[207],"On by default on macOS",[209,211,213],[210],"CUDA",[212],"NVIDIA GPUs",[214],"-DGGML_CUDA=ON",[216,218,220],[217],"Vulkan",[219],"GPUs with a Vulkan driver",[221],"-DGGML_VULKAN=ON",[223,225,227],[224],"CPU",[226],"x86 (AVX2, AVX512, AMX) and ARM (NEON)",[228],"In the default build",{"type":230,"code":231},"code","# macOS includes Metal by default; for NVIDIA GPUs use cmake -B build -DGGML_CUDA=ON\ncmake -B build\ncmake --build build --config Release\n.\u002Fbuild\u002Fbin\u002Fllama-server -m models\u002FLlama-3.1-8B-Instruct-Q4_K_M.gguf -c 8192 -np 4 --host 127.0.0.1 --port 8080",{"type":137,"content":233},[234],"Three flags carry most decisions. -m names the GGUF file, -c sets the context size in tokens (0 means the value stored in the model), and -np sets the number of parallel slots. The server listens on 127.0.0.1 by default, so other machines cannot reach it until you change --host.",{"type":141,"level":142,"id":119,"text":120},{"type":137,"content":237},[238],"A quant is the number of bits each weight gets, traded against file size and quality. The _K names mark k-quants, and the IQ types are i-quants that reach down to about 2 bits per weight; IQ1_S is 2.00 bits in the README’s table. The README’s example command is a naive Q4_K_M quantisation with default settings, which makes Q4_K_M the sensible place to start.",{"type":192,"head":240,"rows":249},[241,243,245,247],[242],"Quant",[244],"Bits per weight",[246],"Size (GiB)",[248],"Generation (tokens\u002Fs)",[250,259,268,277,286,295],[251,253,255,257],[252],"Q2_K",[254],"3.16",[256],"2.95",[258],"79.85",[260,262,264,266],[261],"Q4_K_M",[263],"4.89",[265],"4.58",[267],"71.93",[269,271,273,275],[270],"Q5_K_M",[272],"5.70",[274],"5.33",[276],"67.23",[278,280,282,284],[279],"Q6_K",[281],"6.56",[283],"6.14",[285],"58.67",[287,289,291,293],[288],"Q8_0",[290],"8.50",[292],"7.95",[294],"50.93",[296,298,300,302],[297],"F16",[299],"16.00",[301],"14.96",[303],"29.17",{"type":137,"content":305},[306,307,312],"Read the table as a trade-off. Size falls with the bit count, and generation gets faster as the file shrinks, because each token has to read the weights again. F16 generates 29.17 tokens a second against 71.93 for Q4_K_M, and needs more than three times the memory. The README measures Llama 3.1 8B but does not say which machine produced the numbers, so read them as relative. The table has no quality column, since the README reports no perplexity or KL divergence. I would run the same evaluation set on each candidate before choosing, as I describe in my post on ",{"tag":308,"to":309,"children":310},"link","\u002Fblog\u002Fllm-evals-for-product-features",[311],"evals for LLM product features",".",{"type":314,"variant":315,"title":316,"body":317},"callout","tip","A default that holds up",[318],[319],"Start with Q4_K_M. Move to Q5_K_M or Q6_K only if your evals show a gap that matters, and drop to Q2_K only when memory forces the choice, with the quality loss measured rather than assumed.",{"type":137,"content":321},[322],"On Apple chips, speed follows memory bandwidth. In the Apple Silicon thread, LLaMA 7B at Q4_0 generates 83.06 tokens a second on the M4 Max, which has 546 GB\u002Fs of memory bandwidth, and 36.41 on the M1 Pro, which has 200 GB\u002Fs. The M2 Ultra reaches 94.27 at 800 GB\u002Fs. The relation is not a straight line: the M1 Pro has a quarter of the M2 Ultra’s bandwidth and gets less than half its speed.",{"type":137,"content":324},[325],"The CUDA thread collects llama-bench results on Llama 2 7B at Q4_0: 186.21 tokens a second for an RTX 4090, 267.81 for an H100 80 GB and 290.02 for an RTX 5090. For planning, one user on a high-end Apple chip gets tens of tokens a second on a 7B model, and a data-centre GPU is roughly three times faster in the same kind of test. The models and builds differ between the threads, so compare orders of magnitude. Neither thread tests a server under concurrent load, which is where a GPU server earns its price.",{"type":314,"variant":327,"title":328,"body":329},"note","Benchmark the build you run",[330],[331],"Builds change quickly. The releases page showed four builds on 9 and 10 October 2026 alone. Run the benchmark for your model, quant and hardware, and record the build number next to the result.",{"type":141,"level":142,"id":13,"text":122},{"type":137,"content":334},[335],"llama-server is where most applications meet llama.cpp. Its OpenAI-compatible routes sit under \u002Fv1: chat completions, completions, responses, models and embeddings. Native routes cover tokenising, template rendering, slot state and health. A Prometheus metrics endpoint exists, but it stays off until you pass --metrics, which is worth knowing before you publish a port.",{"type":137,"content":337},[338],"Parallel slots work like this. With -np on its default of auto, the slots share one KV-cache buffer, which the README switches on in that mode, so a long request can use memory an idle slot would otherwise hold. Continuous batching is on by default too. The README documents the switch (--kv-unified) and a per-slot limit (--kv-unified-per-slot), but it does not spell out how the budget divides when you set the slot count by hand, so measure memory before you size a multi-user server.",{"type":137,"content":340},[341],"Structured output has two routes. --json-schema constrains generation to a JSON schema, and --grammar takes a BNF-like grammar; the README describes both as ways to constrain generations. The chat endpoint also accepts a response_format of type json_schema, as in this request.",{"type":230,"code":343},"curl -s http:\u002F\u002F127.0.0.1:8080\u002Fv1\u002Fchat\u002Fcompletions \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\"messages\": [{\"role\": \"user\", \"content\": \"Extract the author from: Written by Balázs Csorba.\"}], \"response_format\": {\"type\": \"json_schema\", \"schema\": {\"type\": \"object\", \"properties\": {\"author\": {\"type\": \"string\"}}, \"required\": [\"author\"]}}}'",{"type":141,"level":142,"id":124,"text":125},{"type":192,"head":346,"rows":353},[347,349,351],[348],"Cost item",[350],"What you pay",[352],"Note",[354,360,367,374],[355,356,358],[9],[357],"Nothing",[359],"MIT licence, no usage fee",[361,363,365],[362],"Your hardware",[364],"Purchase and electricity",[366],"Memory sets the largest model and quant you can load",[368,370,372],[369],"Rented GPU server",[371],"The host’s price",[373],"Pick an EU region and sign a data processing agreement",[375,377,379],[376],"Model weights",[378],"Each model’s own terms",[380],"The model card states the licence",{"type":137,"content":382},[383],"Data protection is the strongest argument for local inference. If the model runs on hardware you own, prompts, retrieved documents and answers stay on that machine, and no model vendor receives them, so the inference step has no processor. GDPR Article 28 says processing by a processor must be governed by a binding contract that limits the processor to documented instructions. A rented GPU host that handles personal data for you is a processor wherever it sits in the EU, so it needs that contract.",{"type":137,"content":385},[386,387,391],"Local does not mean compliant by default. Article 32 asks controllers and processors for appropriate technical and organisational measures, and lists encryption and pseudonymisation among the examples. On a llama.cpp server the settings that matter are the bind address (127.0.0.1 unless you change --host), the API key (--api-key accepts one or more keys), the metrics endpoint (off unless you pass --metrics) and your own logs, which need the same retention rules as any prompt store. The ",{"tag":308,"to":388,"children":389},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[390],"GDPR and LLM data residency"," notes cover the rest of the checklist.",{"type":314,"variant":393,"title":394,"body":395},"warn","The Docker example binds every interface",[396],[397],"The README’s Docker example starts the server with --host 0.0.0.0. Inside a container that is normal, because the port mapping decides who can reach it. On a laptop it exposes the server to the whole network, so keep 127.0.0.1 or add --api-key first.",{"type":141,"level":142,"id":127,"text":128},{"type":147,"ordered":148,"items":400},[401,406,411,416,421],[402,405],{"tag":152,"children":403},[404],"It is a runtime, not a platform."," You choose the file, the quant, the context size and the flags, and you update the binary yourself. The pace is part of the deal: the releases page showed four builds on 9 and 10 October 2026.",[407,410],{"tag":152,"children":408},[409],"No model management."," You download the GGUF file yourself. The -hf flag fetches a Hugging Face repository and defaults to Q4_K_M, or to the first file in the repo when that quant is missing.",[412,415],{"tag":152,"children":413},[414],"No quality measurement."," The quantisation table gives size and speed only, so the quality check is yours.",[417,420],{"tag":152,"children":418},[419],"Concurrency is a memory question."," Slots and continuous batching exist, but the benchmark threads do not test a server under concurrent load, and the README does not spell out the memory split for hand-set slot counts.",[422,425],{"tag":152,"children":423},[424],"GGUF in vLLM is not a serving shortcut."," vLLM calls its GGUF support highly experimental and under-optimised, and says that for now GGUF is mainly a way to reduce the memory footprint.",{"type":141,"level":142,"id":130,"text":131},{"type":137,"content":428},[429],"Take llama.cpp when you want the engine itself, with the file, the quant and the flags under your control, and when the data should never leave the machine. It is the right base for your own tooling and for a small private server. It is not the first tool for someone who wants one command to pull a model, and it is not the serving layer for a busy GPU service. If you are choosing between local options, LM Studio runs llama.cpp on Mac, Windows and Linux and uses MLX on Apple Silicon, so it suits people who want an app with a user interface.",{"type":147,"ordered":431,"items":432},true,[433,438,442,452],[434,437],{"tag":152,"children":435},[436],"Adopt it if"," you want to choose the GGUF file, the quant and every flag on hardware you control.",[439,441],{"tag":152,"children":440},[436]," you want a server you can reason about line by line, with every setting visible in your own command.",[443,446,447,451],{"tag":152,"children":444},[445],"Do not adopt it if"," you want one command that downloads and manages models. Use ",{"tag":308,"to":448,"children":449},"\u002Ftools\u002Follama",[450],"Ollama",", which lists llama.cpp among its supported backends.",[453,456,457,461],{"tag":152,"children":454},[455],"Do not adopt it for a shared GPU service"," with many users at once. Benchmark ",{"tag":308,"to":458,"children":459},"\u002Ftools\u002Fvllm",[460],"vLLM"," first, because its core is PagedAttention and continuous batching.",{"type":141,"level":142,"id":133,"text":134},{"type":147,"ordered":148,"items":464},[465,469,472,475,478,481,484,487,490,493,496,499,502,505],[466],{"tag":467,"href":24,"children":468},"a",[36],[470],{"tag":467,"href":39,"children":471},[38],[473],{"tag":467,"href":42,"children":474},[41],[476],{"tag":467,"href":45,"children":477},[44],[479],{"tag":467,"href":48,"children":480},[47],[482],{"tag":467,"href":51,"children":483},[50],[485],{"tag":467,"href":54,"children":486},[53],[488],{"tag":467,"href":57,"children":489},[56],[491],{"tag":467,"href":60,"children":492},[59],[494],{"tag":467,"href":63,"children":495},[62],[497],{"tag":467,"href":66,"children":498},[65],[500],{"tag":467,"href":69,"children":501},[68],[503],{"tag":467,"href":72,"children":504},[71],[506],{"tag":467,"href":75,"children":507},[74],[509,600,713,819],{"slug":510,"published":511,"minutes":6,"category":7,"tags":512,"keywords":518,"about":526,"sources":537,"cover":592,"og":593,"expertise":78,"locales":594,"lang":80,"title":595,"description":596,"coverAlt":597,"url":529,"pricing":598,"kind":599},"deepeval","2026-10-06",[513,514,515,516,517],"LLM evaluation","pytest","LLM-as-a-judge","Red teaming","Open source",[510,519,520,521,522,523,524,525],"deepeval vs ragas","llm evaluation framework","pytest for llm outputs","g-eval metric","llm judge cost","deepeval pricing","llm red teaming open source",[527,530,533,536],{"name":528,"url":529},"DeepEval","https:\u002F\u002Fdeepeval.com",{"name":531,"url":532},"Confident AI","https:\u002F\u002Fwww.confident-ai.com",{"name":534,"url":535},"Pytest","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPytest",{"name":26,"url":27},[538,541,544,547,550,553,556,559,562,565,568,571,574,577,580,583,586,589],{"title":539,"url":540},"DeepEval documentation: getting started","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fgetting-started",{"title":542,"url":543},"GitHub: confident-ai\u002Fdeepeval, the README and licence","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepeval",{"title":545,"url":546},"PyPI: deepeval, the latest release and Python requirement","https:\u002F\u002Fpypi.org\u002Fproject\u002Fdeepeval\u002F",{"title":548,"url":549},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":551,"url":552},"DeepEval documentation: G-Eval","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-llm-evals",{"title":554,"url":555},"DeepEval documentation: Faithfulness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-faithfulness",{"title":557,"url":558},"DeepEval documentation: Tool Correctness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-tool-correctness",{"title":560,"url":561},"DeepEval documentation: generate goldens from documents","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fsynthesizer-generate-from-docs",{"title":563,"url":564},"DeepEval documentation: unit testing in CI\u002FCD","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fevaluation-unit-testing-in-ci-cd",{"title":566,"url":567},"DeepEval FAQ","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Ffaq",{"title":569,"url":570},"DeepEval documentation: data privacy","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fdata-privacy",{"title":572,"url":573},"GitHub: confident-ai\u002Fdeepteam, the red-teaming framework","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam",{"title":575,"url":576},"Confident AI pricing","https:\u002F\u002Fwww.confident-ai.com\u002Fpricing",{"title":578,"url":579},"Confident AI documentation: data residency","https:\u002F\u002Fwww.confident-ai.com\u002Fdocs\u002Fsettings\u002Fdata-residency",{"title":581,"url":582},"Confident AI subprocessor list","https:\u002F\u002Fwww.confident-ai.com\u002Fsubprocessors-list",{"title":584,"url":585},"GitHub: explodinggradients\u002Fragas, the README","https:\u002F\u002Fgithub.com\u002Fexplodinggradients\u002Fragas",{"title":587,"url":588},"GitHub: promptfoo\u002Fpromptfoo, the README","https:\u002F\u002Fgithub.com\u002Fpromptfoo\u002Fpromptfoo",{"title":590,"url":591},"Braintrust pricing","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fog.jpg",[80,81,82],"DeepEval review: pytest for LLM outputs, and the judge bill","DeepEval runs LLM checks as pytest-style tests, with built-in judge metrics. The library is free and Apache-2.0; the judge calls and the data flow are the real cost.","Cover art for the DeepEval review: a test case passed through a judge metric to a pass or fail gate","Apache-2.0 · free; Confident AI from $0, Starter $200 a month","Evaluation framework",{"slug":601,"published":511,"minutes":6,"category":7,"tags":602,"keywords":608,"about":616,"sources":629,"cover":705,"og":706,"expertise":78,"locales":707,"lang":80,"title":708,"description":709,"coverAlt":710,"url":619,"pricing":711,"kind":712},"dspy",[603,604,605,606,607],"Prompt optimisation","LLM programs","MIPROv2","GEPA","Python",[601,609,610,611,612,613,614,615],"dspy tutorial","dspy optimizer","miprov2","gepa prompt optimization","dspy vs prompt engineering","automatic prompt optimization","dspy signatures",[617,620,623,626],{"name":618,"url":619},"DSPy","https:\u002F\u002Fdspy.ai",{"name":621,"url":622},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",{"name":624,"url":625},"Bayesian optimization","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBayesian_optimization",{"name":627,"url":628},"Fine-tuning (deep learning)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FFine-tuning_(deep_learning)",[630,633,636,639,642,645,648,651,654,657,660,663,666,669,672,675,678,681,684,687,690,693,696,699,702],{"title":631,"url":632},"DSPy home and cost example","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002F",{"title":634,"url":635},"DSPy installation","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Finstallation\u002F",{"title":637,"url":638},"DSPy: program, don't prompt","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fprogram-dont-prompt\u002F",{"title":640,"url":641},"DSPy signatures","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fdiving-deeper\u002Fsignatures-in-depth\u002F",{"title":643,"url":644},"DSPy metrics","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fmetrics\u002F",{"title":646,"url":647},"DSPy Example API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fprimitives\u002FExample\u002F",{"title":649,"url":650},"DSPy Predict API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FPredict\u002F",{"title":652,"url":653},"DSPy ChainOfThought API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FChainOfThought\u002F",{"title":655,"url":656},"DSPy ReAct API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FReAct\u002F",{"title":658,"url":659},"DSPy optimizers index","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002F",{"title":661,"url":662},"DSPy MIPROv2 API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FMIPROv2\u002F",{"title":664,"url":665},"MIPROv2 source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fmipro_optimizer_v2.py",{"title":667,"url":668},"DSPy BootstrapFewShot API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFewShot\u002F",{"title":670,"url":671},"DSPy BootstrapFinetune API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFinetune\u002F",{"title":673,"url":674},"DSPy GEPA guide","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fgepa-optimization\u002F",{"title":676,"url":677},"GEPA source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fgepa\u002Fgepa.py",{"title":679,"url":680},"DSPy saving programs","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fsaving\u002F",{"title":682,"url":683},"DSPy caching","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fcache\u002F",{"title":685,"url":686},"DSPy LM API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodels\u002FLM\u002F",{"title":688,"url":689},"DSPy FAQ","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ffaqs\u002F",{"title":691,"url":692},"DSPy GitHub repository","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy",{"title":694,"url":695},"DSPy 3.4.0 release","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Freleases\u002Ftag\u002F3.4.0",{"title":697,"url":698},"DSPy paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.03714",{"title":700,"url":701},"MIPRO paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.11695",{"title":703,"url":704},"GEPA paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2507.19457","\u002Fimages\u002Fblog\u002Fdspy\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdspy\u002Fog.jpg",[80,81,82],"DSPy review: compile your prompts against a metric, not by hand","DSPy compiles prompts from signatures, a metric and examples. What the optimisers cost in model calls, when they pay off, and when a hand-written prompt wins.","Cover art for the DSPy review: a signature and a metric feed an optimiser that compiles a saved program.","MIT · free, you pay the model API","Prompt optimisation framework",{"slug":714,"published":5,"minutes":715,"category":7,"tags":716,"keywords":722,"about":730,"sources":742,"cover":812,"og":813,"expertise":78,"locales":814,"lang":80,"title":815,"description":816,"coverAlt":817,"url":754,"pricing":818,"kind":717},"opik",7,[717,718,719,720,721],"LLM observability","Tracing","Evaluation","OpenTelemetry","Self-hosting",[714,723,724,725,726,727,728,729],"opik review","opik vs langfuse","opik pricing","self-hosted llm tracing","opik opentelemetry","llm as a judge metrics","opik guardrails",[731,734,736,739],{"name":732,"url":733},"Comet ML","https:\u002F\u002Fwww.comet.com",{"name":720,"url":735},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":737,"url":738},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",{"name":740,"url":741},"Apache License","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_License",[743,746,749,752,755,758,761,764,767,770,773,776,779,782,785,788,791,794,797,800,803,806,809],{"title":744,"url":745},"Opik repository README on GitHub","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik",{"title":747,"url":748},"Opik LICENSE file: Apache License 2.0","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik\u002Fblob\u002Fmain\u002FLICENSE",{"title":750,"url":751},"Opik on PyPI: package metadata","https:\u002F\u002Fpypi.org\u002Fpypi\u002Fopik\u002Fjson",{"title":753,"url":754},"Opik product page","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fproducts\u002Fopik\u002F",{"title":756,"url":757},"Opik pricing","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fpricing\u002F",{"title":759,"url":760},"Opik docs: run locally with Docker Compose","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Flocal_deployment",{"title":762,"url":763},"Opik docs: Kubernetes deployment with Helm","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Fkubernetes",{"title":765,"url":766},"Opik docs: platform architecture","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Farchitecture",{"title":768,"url":769},"Opik docs: tracing getting started","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Ftracing\u002Fgetting-started",{"title":771,"url":772},"Opik docs: OpenTelemetry integration","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fintegrations\u002Fopentelemetry",{"title":774,"url":775},"Opik docs: evaluation metrics overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fevaluation\u002Fmetrics\u002Foverview",{"title":777,"url":778},"Opik docs: online evaluation rules","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fonline-evaluation\u002Frules",{"title":780,"url":781},"Opik docs: Prompt Library overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Fprompt-library\u002Foverview",{"title":783,"url":784},"Opik docs: optimisation algorithms","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Foptimization-runs\u002Falgorithms\u002Foverview",{"title":786,"url":787},"Opik docs: guardrails overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Foverview",{"title":789,"url":790},"Opik docs: guardrails server","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Fserver",{"title":792,"url":793},"Opik docs: SDK anonymizers","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fgateway-guardrails\u002Fanonymizers",{"title":795,"url":796},"Opik docs: data anonymization","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fadministration\u002Fdata_anonymization",{"title":798,"url":799},"Comet ML privacy policy","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fprivacy-policy\u002F",{"title":801,"url":802},"Comet Trust Center","https:\u002F\u002Ftrust.comet.com\u002F",{"title":804,"url":805},"Langfuse repository README (licence)","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse",{"title":807,"url":808},"Arize Phoenix repository README (licence)","https:\u002F\u002Fgithub.com\u002FArize-ai\u002Fphoenix",{"title":810,"url":811},"LangChain pricing (LangSmith plans)","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fopik\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopik\u002Fog.jpg",[80,81,82],"Opik review: open-source tracing and evals, with a US-hosted cloud","Opik puts traces, LLM-as-a-judge metrics, prompt versions and an optimiser on one Apache 2.0 platform. Free to self-host, Pro cloud at $19 a month, US-hosted.","Cover art for the Opik review: a trace moves from the app through the backend to a judge and a score gate.","Apache 2.0 · free to self-host · Pro cloud from $19 a month",{"slug":820,"published":821,"minutes":822,"category":7,"tags":823,"keywords":826,"about":833,"sources":837,"cover":855,"og":856,"expertise":78,"locales":857,"lang":80,"title":858,"description":859,"coverAlt":860,"url":861,"pricing":862,"kind":87},"ollama","2026-09-29",11,[12,824,9,10,825],"Open models","Model serving",[820,827,828,829,830,831,832],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[834],{"name":835,"url":836},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[838,841,843,846,849,852],{"title":839,"url":840},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":842,"url":60},"Ollama on GitHub, with the MIT LICENSE file",{"title":844,"url":845},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":847,"url":848},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":850,"url":851},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":853,"url":854},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[80,81,82],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Cover art for the Ollama review: a request on port 11434 passes the scheduler and the engine and returns streamed tokens, with model loading and idle unloading noted below.","https:\u002F\u002Follama.com","MIT · free for personal use",1791636874674]