[{"data":1,"prerenderedAt":853},["ShallowReactive",2],{"tool-dspy-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":35,"cover":111,"og":112,"expertise":113,"locales":114,"lang":115,"title":118,"description":119,"coverAlt":120,"url":25,"pricing":121,"kind":122,"metaTitle":123,"takeaways":124,"faq":130,"toc":143,"blocks":171,"others":534},"dspy","2026-10-06",8,"llmops",[9,10,11,12,13],"Prompt optimisation","LLM programs","MIPROv2","GEPA","Python",[4,15,16,17,18,19,20,21],"dspy tutorial","dspy optimizer","miprov2","gepa prompt optimization","dspy vs prompt engineering","automatic prompt optimization","dspy signatures",[23,26,29,32],{"name":24,"url":25},"DSPy","https:\u002F\u002Fdspy.ai",{"name":27,"url":28},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",{"name":30,"url":31},"Bayesian optimization","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBayesian_optimization",{"name":33,"url":34},"Fine-tuning (deep learning)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FFine-tuning_(deep_learning)",[36,39,42,45,48,51,54,57,60,63,66,69,72,75,78,81,84,87,90,93,96,99,102,105,108],{"title":37,"url":38},"DSPy home and cost example","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002F",{"title":40,"url":41},"DSPy installation","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Finstallation\u002F",{"title":43,"url":44},"DSPy: program, don't prompt","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fprogram-dont-prompt\u002F",{"title":46,"url":47},"DSPy signatures","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fdiving-deeper\u002Fsignatures-in-depth\u002F",{"title":49,"url":50},"DSPy metrics","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fmetrics\u002F",{"title":52,"url":53},"DSPy Example API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fprimitives\u002FExample\u002F",{"title":55,"url":56},"DSPy Predict API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FPredict\u002F",{"title":58,"url":59},"DSPy ChainOfThought API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FChainOfThought\u002F",{"title":61,"url":62},"DSPy ReAct API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodules\u002FReAct\u002F",{"title":64,"url":65},"DSPy optimizers index","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002F",{"title":67,"url":68},"DSPy MIPROv2 API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FMIPROv2\u002F",{"title":70,"url":71},"MIPROv2 source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fmipro_optimizer_v2.py",{"title":73,"url":74},"DSPy BootstrapFewShot API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFewShot\u002F",{"title":76,"url":77},"DSPy BootstrapFinetune API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Foptimizers\u002FBootstrapFinetune\u002F",{"title":79,"url":80},"DSPy GEPA guide","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fgetting-started\u002Fgepa-optimization\u002F",{"title":82,"url":83},"GEPA source code","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Fblob\u002Fmain\u002Fdspy\u002Fteleprompt\u002Fgepa\u002Fgepa.py",{"title":85,"url":86},"DSPy saving programs","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fsaving\u002F",{"title":88,"url":89},"DSPy caching","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ftutorials\u002Fcache\u002F",{"title":91,"url":92},"DSPy LM API","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Fapi\u002Fmodels\u002FLM\u002F",{"title":94,"url":95},"DSPy FAQ","https:\u002F\u002Fdspy.ai\u002Fcurrent\u002Ffaqs\u002F",{"title":97,"url":98},"DSPy GitHub repository","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy",{"title":100,"url":101},"DSPy 3.4.0 release","https:\u002F\u002Fgithub.com\u002Fstanfordnlp\u002Fdspy\u002Freleases\u002Ftag\u002F3.4.0",{"title":103,"url":104},"DSPy paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.03714",{"title":106,"url":107},"MIPRO paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.11695",{"title":109,"url":110},"GEPA paper, arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2507.19457","\u002Fimages\u002Fblog\u002Fdspy\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdspy\u002Fog.jpg","ai-engineer",[115,116,117],"en","de","hu","DSPy review: compile your prompts against a metric, not by hand","DSPy compiles prompts from signatures, a metric and examples. What the optimisers cost in model calls, when they pay off, and when a hand-written prompt wins.","Cover art for the DSPy review: a signature and a metric feed an optimiser that compiles a saved program.","MIT · free, you pay the model API","Prompt optimisation framework","DSPy review: prompts tuned against a metric · Balázs Csorba",[125,126,127,128,129],"DSPy replaces hand-written prompts with signatures, modules and optimisers that search for better prompts against a metric you write.","An optimisation run is paid for in model calls. On MIPROv2's light setting, a one-predictor program makes roughly 750 of them before bootstrapping.","Write the metric first. An optimiser can only improve what the metric measures, and that function is where most of the work sits.","A compiled program is a JSON file of instructions and demonstrations, so it needs version control, and its demonstrations are your training data.","For one prompt on a stable model, a reviewed prompt and an eval suite are simpler. DSPy earns its place in multi-call pipelines that change.",[131,134,137,140],{"q":132,"a":133},"Is DSPy free to use?","The library is MIT-licensed and free. You pay for every model call, including the calls the optimiser makes while it searches. The docs' own example runs cost a few dollars.",{"q":135,"a":136},"How many examples does DSPy need?","I found no fixed minimum in the docs. The FAQ asks for a few example inputs, with labels only when the metric needs them. MIPROv2's light setting uses at most 100 examples for validation.",{"q":138,"a":139},"Can I point DSPy at a local model?","dspy.LM takes LiteLLM-style provider and model strings, so a local endpoint may work if LiteLLM can call it. I did not test one for this review, so check it on your own setup first.",{"q":141,"a":142},"Do I need to re-optimise when I change models?","Yes. The FAQ names a change of target LM as a reason to recompile, because the compiler maps the program onto new prompts. Keep the previous compiled file so you can compare and roll back.",[144,147,150,153,156,159,162,165,168],{"id":145,"title":146},"what-it-is","What it is",{"id":148,"title":149},"how-it-works","How it works",{"id":151,"title":152},"getting-started","Getting started",{"id":154,"title":155},"signatures-and-modules","Signatures and modules",{"id":157,"title":158},"optimisers","Optimisers and what they need",{"id":160,"title":161},"cost-and-deployment","Cost, deployment and data",{"id":163,"title":164},"where-it-falls-short","Where it falls short",{"id":166,"title":167},"verdict","Verdict",{"id":169,"title":170},"sources","Sources",[172,176,179,188,218,219,222,231,232,234,237,238,241,258,261,262,265,311,314,331,338,341,344,347,348,385,388,391,398,404,405,408,411,414,417,418,421,442,455,456],{"type":173,"content":174},"paragraph",[175],"DSPy is an open-source Python framework that treats an LLM pipeline as code. You declare inputs and outputs, pick a module, and let an optimiser search for the instructions and examples that score best on a metric you wrote. The verdict up front: use it for a multi-call pipeline you can measure and expect to change. Skip it for one prompt on a stable model – a reviewed prompt and an eval suite are simpler.",{"type":177,"level":178,"id":145,"text":146},"heading",2,{"type":173,"content":180},[181,182,187],"DSPy calls itself the framework for programming, not prompting, language models. Signatures declare inputs and outputs. Modules decide how the model is asked, from a plain prediction to step-by-step reasoning or a tool loop. Optimisers compile a program against a metric by changing its instructions and examples. It sits between prompt engineering and fine-tuning: the prompt becomes an artefact the optimiser writes. For the wider choice, read ",{"tag":183,"to":184,"children":185},"link","\u002Fblog\u002Ffine-tuning-vs-rag-vs-prompting",[186],"my decision guide on prompting, retrieval and fine-tuning",".",{"type":189,"ordered":190,"items":191},"list",false,[192,203,208,213],[193,197,198,202],{"tag":194,"children":195},"strong",[196],"MIT licence, free to use. ","It installs with ",{"tag":199,"children":200},"code",[201],"pip install dspy"," and needs Python 3.10 or newer.",[204,207],{"tag":194,"children":205},[206],"Latest release 3.4.0, ","published on 25 September. The repository showed 38.6k stars when I checked.",[209,212],{"tag":194,"children":210},[211],"14 optimisers in the API reference, ","from BootstrapFewShot to MIPROv2, GEPA and BootstrapFinetune.",[214,217],{"tag":194,"children":215},[216],"No hosted service that I could find. ","It runs in your process and calls the provider you configure.",{"type":177,"level":178,"id":148,"text":149},{"type":173,"content":220},[221],"A module is built from a signature. At run time the adapter turns the signature into the system message, the model answers, and DSPy parses the output fields. Without an optimiser, the program is only as good as its wording. The optimiser adds a metric, a Python function that scores one prediction, usually from 0.0 to 1.0, and example inputs. It runs the program over the examples, proposes new instructions and demonstrations, keeps the best combination and returns a compiled program.",{"type":223,"attrs":224,"inner":228,"caption":229},"diagram",{"viewBox":225,"role":226,"aria-labelledby":227},"0 0 720 330","img","d1-dy-t d1-dy-d","\u003Ctitle id=\"d1-dy-t\">How DSPy compiles a program\u003C\u002Ftitle>\u003Cdesc id=\"d1-dy-d\">A signature and a module form the program. A metric and a set of examples feed the optimiser, which searches instructions and demonstrations and returns a compiled program. The compiled program can be saved as a JSON file and called in the application in place of the original module.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"d1-ah\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">Compile a program\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">same program, better prompts\u003C\u002Ftext>\u003Crect x=\"380\" y=\"52\" width=\"140\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"450\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Metric\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"102\" text-anchor=\"middle\" class=\"d-small\">scores one answer\u003C\u002Ftext>\u003Crect x=\"20\" y=\"170\" width=\"140\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"90\" y=\"198\" text-anchor=\"middle\" class=\"d-text\">Signature\u003C\u002Ftext>\u003Ctext x=\"90\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">inputs and outputs\u003C\u002Ftext>\u003Crect x=\"200\" y=\"170\" width=\"140\" height=\"62\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"270\" y=\"198\" text-anchor=\"middle\" class=\"d-text\">Module\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">Predict, CoT, ReAct\u003C\u002Ftext>\u003Crect x=\"380\" y=\"170\" width=\"140\" height=\"62\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"450\" y=\"198\" text-anchor=\"middle\" class=\"d-text\">Optimiser\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">MIPROv2, GEPA\u003C\u002Ftext>\u003Crect x=\"560\" y=\"170\" width=\"140\" height=\"62\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"630\" y=\"198\" text-anchor=\"middle\" class=\"d-text\">Compiled program\u003C\u002Ftext>\u003Ctext x=\"630\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">saved as JSON\u003C\u002Ftext>\u003Crect x=\"380\" y=\"262\" width=\"140\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"450\" y=\"286\" text-anchor=\"middle\" class=\"d-text\">Examples\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"306\" text-anchor=\"middle\" class=\"d-small\">train and validation\u003C\u002Ftext>\u003Cpath d=\"M160 201 H198\" class=\"d-line\" marker-end=\"url(#d1-ah)\" \u002F>\u003Cpath d=\"M340 201 H378\" class=\"d-line\" marker-end=\"url(#d1-ah)\" \u002F>\u003Cpath d=\"M520 201 H558\" class=\"d-line\" marker-end=\"url(#d1-ah)\" \u002F>\u003Cpath d=\"M450 114 V168\" class=\"d-line\" marker-end=\"url(#d1-ah)\" \u002F>\u003Cpath d=\"M450 260 V234\" class=\"d-line\" marker-end=\"url(#d1-ah)\" \u002F>\u003Ctext x=\"20\" y=\"282\" class=\"d-small\">The optimiser runs the program on the examples,\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"300\" class=\"d-small\">scores each answer, and keeps the best combination.\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"318\" class=\"d-small\">Then the compiled program replaces the module.\u003C\u002Ftext>",[230],"The metric and examples feed the optimiser, which compiles the signature and module into a program you can save.",{"type":177,"level":178,"id":151,"text":152},{"type":199,"code":233},"import os\nimport dspy\n\nlm = dspy.LM(\"openai\u002Fgpt-5-nano\", api_key=os.environ[\"OPENAI_API_KEY\"])\ndspy.configure(lm=lm)\n\n\nclass ExtractIntent(dspy.Signature):\n    \"\"\"Classify the customer's intent in one short email.\"\"\"\n\n    email: str = dspy.InputField()\n    intent: str = dspy.OutputField(desc=\"one of: order, return, invoice, other\")\n\n\nextract = dspy.ChainOfThought(ExtractIntent)\n\n\ndef intent_metric(example, prediction, trace=None):\n    return float(prediction.intent.strip().lower() == example.intent.strip().lower())\n\n\ntrainset = [\n    dspy.Example(email=\"Where is my parcel 4411?\", intent=\"order\").with_inputs(\"email\"),\n    dspy.Example(email=\"I want my money back for the shoes.\", intent=\"return\").with_inputs(\"email\"),\n    # more labelled emails from your own inbox\n]\n\noptimizer = dspy.MIPROv2(metric=intent_metric, auto=\"light\")\ncompiled = optimizer.compile(extract, trainset=trainset)\ncompiled.save(\"intent_v1.json\")",{"type":173,"content":235},[236],"The metric compares predicted and labelled intents, so this example needs labels. The compile call is where the cost sits: light makes hundreds of model calls, so run it on a small set, check the saved file, then scale up.",{"type":177,"level":178,"id":154,"text":155},{"type":173,"content":239},[240],"A signature is the contract. The string form, such as question -> answer, is shorthand. The class form adds a docstring, which becomes the instruction, and typed fields. Field order matters, because reordering inputs or outputs changes the prompt. Test every signature edit as a prompt change.",{"type":189,"ordered":190,"items":242},[243,248,253],[244,247],{"tag":194,"children":245},[246],"dspy.Predict ","maps inputs to outputs with a language model. Its keyword arguments go to that model.",[249,252],{"tag":194,"children":250},[251],"dspy.ChainOfThought ","reasons step by step first. It adds a reasoning field you can customise.",[254,257],{"tag":194,"children":255},[256],"dspy.ReAct ","runs a reason-and-act loop over tools. Its max_iters defaults to 20.",{"type":173,"content":259},[260],"The homepage sums this up as 'same interface, different strategy'. The strategy also sets the bill: reasoning fields add output tokens to every call, and each ReAct step is another model call.",{"type":177,"level":178,"id":157,"text":158},{"type":173,"content":263},[264],"DSPy ships 14 optimisers, from BootstrapFewShot, which collects demonstrations, to GEPA, which rewrites instructions from the metric's feedback. All of them need a metric. The FAQ asks for a task, a metric and a few example inputs, with labels only where the metric needs them. The table covers the four I looked at most closely.",{"type":266,"head":267,"rows":276},"table",[268,270,272,274],[269],"Optimiser",[271],"Changes",[273],"Needs",[275],"Main cost",[277,286,294,302],[278,280,282,284],[279],"BootstrapFewShot",[281],"Few-shot demonstrations",[283],"Metric and training examples",[285],"Program runs, one attempt per example",[287,288,290,292],[11],[289],"Instructions and demonstrations",[291],"Metric and training set",[293],"Trials of 35 examples, plus full validation passes",[295,296,298,300],[12],[297],"Instructions, rewritten by a reflection model",[299],"Metric with feedback and a reflection model",[301],"Budget set by validation size and predictor count",[303,305,307,309],[304],"BootstrapFinetune",[306],"Fine-tuned model per predictor",[308],"Traces and a model you can fine-tune",[310],"One job per model, or per predictor",{"type":173,"content":312},[313],"MIPROv2 is the usual starting point, so its arithmetic matters. Its auto setting fixes the search. For a one-predictor few-shot program, light runs about ten trials and validates on at most 100 examples, medium runs 18 trials on 300 and heavy runs 27 on 1,000. Each trial scores a 35-example minibatch, and a full validation pass runs every sixth trial, at the last trial and for the unoptimised program.",{"type":189,"ordered":190,"items":315},[316,321,326],[317,320],{"tag":194,"children":318},[319],"light: about 750 runs, ","350 for trials and 400 for full passes.",[322,325],{"tag":194,"children":323},[324],"medium: about 2,100 runs, ","630 for trials and 1,500 for full passes.",[327,330],{"tag":194,"children":328},[329],"heavy: about 7,900 runs, ","945 for trials and 7,000 for full passes.",{"type":332,"variant":333,"title":334,"body":335},"callout","note","How to read the counts",[336],[337],"My counts come from the optimiser code, for one predictor with few-shot demonstrations. Bootstrapping and proposal calls come on top, and each extra predictor adds trials and calls. Treat them as orders of magnitude, and count the calls in your first run.",{"type":173,"content":339},[340],"Tokens follow the calls. The FAQ, flagged as possibly out of date for DSPy 2.5 and 2.6, reports about six minutes, 3,200 calls, 2.7 million input tokens and 156,000 output tokens, for about $3 at the OpenAI pricing of the time. That is roughly 850 input and 50 output tokens per call, so the 750 calls of a light run come to about 0.6 million input tokens and 40,000 output tokens. The homepage's current example, GEPA with auto set to medium on 200 examples and gpt-5.4-mini, is listed at $2.18.",{"type":173,"content":342},[343],"GEPA spends its budget differently. Its metric returns a score and feedback text, and a reflection model reads the examples and their scores and proposes rewritten instructions. The guide recommends a larger reflection model than the one you optimise, and the constructor requires one unless you pass a custom proposer. Light targets about six candidate prompts, and the code turns that into a metric-call budget with several full validation passes inside it.",{"type":173,"content":345},[346],"The papers make the strongest case, each on its authors' own tasks. The DSPy paper reports pipelines beating standard few-shot prompting by over 25% and 65% for GPT-3.5 and llama2-13b-chat respectively. The MIPROv2 paper reports wins on five of seven multi-stage programs, with gains up to 13% accuracy. The GEPA paper reports 6% on average over GRPO, a reinforcement-learning baseline, with up to 35 times fewer rollouts.",{"type":177,"level":178,"id":160,"text":161},{"type":266,"head":349,"rows":356},[350,352,354],[351],"Item",[353],"Price",[355],"What it covers",[357,364,371,378],[358,360,362],[359],"DSPy library",[361],"MIT · free",[363],"Installed with pip, runs in your process",[365,367,369],[366],"Model calls at run time",[368],"Your provider's rate",[370],"Every call the program makes, billed per token",[372,374,376],[373],"Vendor example run",[375],"$2.18",[377],"GEPA, auto medium, 200 examples, gpt-5.4-mini",[379,381,383],[380],"Older FAQ run",[382],"About $3",[384],"3,200 calls, 2.7 million input and 156,000 output tokens",{"type":173,"content":386},[387],"Treat the compiled program as a build artefact. Save it as JSON, which the docs call safer and readable, beside the signatures in the same repository, with a version in the file name and the DSPy version pinned. It holds the signature, the demonstrations and the model for each predictor. Loading needs the same program built in code first.",{"type":173,"content":389},[390],"Model changes trigger a recompile, because the compiler maps the program onto new prompts for the new model. Keep the old file, run the metric on both, and promote the new one only if it wins. Pin the version too: the LM page describes an auto engine that prefers a newer backend, so an upgrade can change behaviour.",{"type":173,"content":392},[393,394,187],"The data path is the one you configure, so the processor and region questions match those for any model API. Three things also keep copies of your data. The LM cache is on by default, in memory and on disk. A saved program carries demonstrations drawn from your training data, so its JSON can hold personal data. A GEPA reflection model reads your examples and their scores, which makes it a second processor. For EU options, see ",{"tag":183,"to":395,"children":396},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[397],"my GDPR article on EU data residency",{"type":332,"variant":399,"title":400,"body":401},"warn","Never load a pickle you did not write",[402],[403],"The docs say a pickle can execute arbitrary code, and save_program=True uses cloudpickle with the same risk. Load only the JSON state, from your own repository.",{"type":177,"level":178,"id":163,"text":164},{"type":173,"content":406},[407],"Most tasks do not need an optimiser. The FAQ concedes that for extremely simple settings a plain prompt might work just fine, and you still write the tools, retries and parsing. If you cannot write a metric that matches what a user would call correct, the optimiser improves whatever you did measure, which is not the same thing.",{"type":173,"content":409},[410],"It beats hand-written prompts most clearly when several calls depend on each other, the model changes often, and the output can be scored. Against fine-tuning it is the cheaper and more reversible option for most teams. BootstrapFinetune compiles the program into fine-tuning jobs, but then you serve and version fine-tuned models, one per model or per predictor.",{"type":173,"content":412},[413],"The bill and the metric are the other weak points. A light run makes hundreds of calls, heavy runs thousands, and the validation set drives most of the price. The vendor line that a small, cheap model can often match or beat a hand-prompted frontier one is a hypothesis to test on your own data.",{"type":173,"content":415},[416],"The API is still moving. The 3.4.0 release notes list a breaking change to rlm(...) and remove the old dspy.LMRequest and dspy.LMResponse exports, so read them before each upgrade.",{"type":177,"level":178,"id":166,"text":167},{"type":173,"content":419},[420],"Adopt DSPy when a pipeline is multi-step, measurable and changing. For one prompt on a model you never change, a reviewed prompt and a regression suite do the job. If you cannot write the metric, do not compile anything yet.",{"type":189,"ordered":422,"items":423},true,[424,429,433,438],[425,428],{"tag":194,"children":426},[427],"Adopt it if ","several model calls must agree and their output can be scored automatically.",[430,432],{"tag":194,"children":431},[427],"the model or the data changes often, since recompiling beats rewriting prompts.",[434,437],{"tag":194,"children":435},[436],"Do not adopt it if ","the task is one prompt on a stable model.",[439,441],{"tag":194,"children":440},[436],"nobody will write and maintain the metric.",{"type":173,"content":443},[444,445,449,450,454],"Three alternatives cover most of the rest. If you want prompts kept in code and changes gated by evals, ",{"tag":183,"to":446,"children":447},"\u002Ftools\u002Fpromptfoo",[448],"Promptfoo"," is the closer fit. If the problem is explicit state and control flow, look at ",{"tag":183,"to":451,"children":452},"\u002Ftools\u002Flanggraph",[453],"LangGraph",", and combine the two if you need both. If the model must learn a format or a style, fine-tuning is the lever, as the decision guide explains.",{"type":177,"level":178,"id":169,"text":170},{"type":189,"ordered":190,"items":457},[458,462,465,468,471,474,477,480,483,486,489,492,495,498,501,504,507,510,513,516,519,522,525,528,531],[459],{"tag":460,"href":38,"children":461},"a",[37],[463],{"tag":460,"href":41,"children":464},[40],[466],{"tag":460,"href":44,"children":467},[43],[469],{"tag":460,"href":47,"children":470},[46],[472],{"tag":460,"href":50,"children":473},[49],[475],{"tag":460,"href":53,"children":476},[52],[478],{"tag":460,"href":56,"children":479},[55],[481],{"tag":460,"href":59,"children":482},[58],[484],{"tag":460,"href":62,"children":485},[61],[487],{"tag":460,"href":65,"children":488},[64],[490],{"tag":460,"href":68,"children":491},[67],[493],{"tag":460,"href":71,"children":494},[70],[496],{"tag":460,"href":74,"children":497},[73],[499],{"tag":460,"href":77,"children":500},[76],[502],{"tag":460,"href":80,"children":503},[79],[505],{"tag":460,"href":83,"children":506},[82],[508],{"tag":460,"href":86,"children":509},[85],[511],{"tag":460,"href":89,"children":512},[88],[514],{"tag":460,"href":92,"children":515},[91],[517],{"tag":460,"href":95,"children":518},[94],[520],{"tag":460,"href":98,"children":521},[97],[523],{"tag":460,"href":101,"children":524},[100],[526],{"tag":460,"href":104,"children":527},[103],[529],{"tag":460,"href":107,"children":530},[106],[532],{"tag":460,"href":110,"children":533},[109],[535,627,703,809],{"slug":536,"published":5,"minutes":6,"category":7,"tags":537,"keywords":543,"about":551,"sources":564,"cover":619,"og":620,"expertise":113,"locales":621,"lang":115,"title":622,"description":623,"coverAlt":624,"url":554,"pricing":625,"kind":626},"deepeval",[538,539,540,541,542],"LLM evaluation","pytest","LLM-as-a-judge","Red teaming","Open source",[536,544,545,546,547,548,549,550],"deepeval vs ragas","llm evaluation framework","pytest for llm outputs","g-eval metric","llm judge cost","deepeval pricing","llm red teaming open source",[552,555,558,561],{"name":553,"url":554},"DeepEval","https:\u002F\u002Fdeepeval.com",{"name":556,"url":557},"Confident AI","https:\u002F\u002Fwww.confident-ai.com",{"name":559,"url":560},"Pytest","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPytest",{"name":562,"url":563},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[565,568,571,574,577,580,583,586,589,592,595,598,601,604,607,610,613,616],{"title":566,"url":567},"DeepEval documentation: getting started","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fgetting-started",{"title":569,"url":570},"GitHub: confident-ai\u002Fdeepeval, the README and licence","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepeval",{"title":572,"url":573},"PyPI: deepeval, the latest release and Python requirement","https:\u002F\u002Fpypi.org\u002Fproject\u002Fdeepeval\u002F",{"title":575,"url":576},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":578,"url":579},"DeepEval documentation: G-Eval","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-llm-evals",{"title":581,"url":582},"DeepEval documentation: Faithfulness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-faithfulness",{"title":584,"url":585},"DeepEval documentation: Tool Correctness","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-tool-correctness",{"title":587,"url":588},"DeepEval documentation: generate goldens from documents","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fsynthesizer-generate-from-docs",{"title":590,"url":591},"DeepEval documentation: unit testing in CI\u002FCD","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fevaluation-unit-testing-in-ci-cd",{"title":593,"url":594},"DeepEval FAQ","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Ffaq",{"title":596,"url":597},"DeepEval documentation: data privacy","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fdata-privacy",{"title":599,"url":600},"GitHub: confident-ai\u002Fdeepteam, the red-teaming framework","https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam",{"title":602,"url":603},"Confident AI pricing","https:\u002F\u002Fwww.confident-ai.com\u002Fpricing",{"title":605,"url":606},"Confident AI documentation: data residency","https:\u002F\u002Fwww.confident-ai.com\u002Fdocs\u002Fsettings\u002Fdata-residency",{"title":608,"url":609},"Confident AI subprocessor list","https:\u002F\u002Fwww.confident-ai.com\u002Fsubprocessors-list",{"title":611,"url":612},"GitHub: explodinggradients\u002Fragas, the README","https:\u002F\u002Fgithub.com\u002Fexplodinggradients\u002Fragas",{"title":614,"url":615},"GitHub: promptfoo\u002Fpromptfoo, the README","https:\u002F\u002Fgithub.com\u002Fpromptfoo\u002Fpromptfoo",{"title":617,"url":618},"Braintrust pricing","https:\u002F\u002Fwww.braintrust.dev\u002Fpricing","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fdeepeval\u002Fog.jpg",[115,116,117],"DeepEval review: pytest for LLM outputs, and the judge bill","DeepEval runs LLM checks as pytest-style tests, with built-in judge metrics. The library is free and Apache-2.0; the judge calls and the data flow are the real cost.","Cover art for the DeepEval review: a test case passed through a judge metric to a pass or fail gate","Apache-2.0 · free; Confident AI from $0, Starter $200 a month","Evaluation framework",{"slug":628,"published":629,"minutes":6,"category":7,"tags":630,"keywords":636,"about":644,"sources":654,"cover":696,"og":697,"expertise":113,"locales":698,"lang":115,"title":699,"description":700,"coverAlt":701,"url":646,"pricing":361,"kind":702},"llama-cpp","2026-10-05",[631,632,633,634,635],"llama.cpp","GGUF","Quantisation","Local inference","llama-server",[631,637,638,639,640,641,642,643],"llama.cpp review","GGUF quantisation levels","llama-server OpenAI compatible","llama.cpp vs Ollama","run LLM locally GDPR","llama.cpp CUDA Metal Vulkan","Q4_K_M vs Q8_0",[645,647,648,651],{"name":631,"url":646},"https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp",{"name":562,"url":563},{"name":649,"url":650},"Quantization (signal processing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FQuantization_(signal_processing)",{"name":652,"url":653},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",[655,657,660,663,666,669,672,675,678,681,684,687,690,693],{"title":656,"url":646},"llama.cpp repository: goals, backends, licence",{"title":658,"url":659},"llama.cpp releases: builds b11538 to b11541","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Freleases",{"title":661,"url":662},"llama.cpp build documentation: CMake flags","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Fdocs\u002Fbuild.md",{"title":664,"url":665},"llama.cpp server README: endpoints, slots and grammars","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fserver\u002FREADME.md",{"title":667,"url":668},"llama.cpp quantisation README: bits, size and speed","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fblob\u002Fmaster\u002Ftools\u002Fquantize\u002FREADME.md",{"title":670,"url":671},"GGUF specification in the ggml repository","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fggml\u002Fblob\u002Fmaster\u002Fdocs\u002Fgguf.md",{"title":673,"url":674},"Performance of llama.cpp on Apple Silicon M-series","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F4167",{"title":676,"url":677},"Performance of llama.cpp on Nvidia CUDA","https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Fdiscussions\u002F15013",{"title":679,"url":680},"Ollama README: supported backends and REST API","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":682,"url":683},"LM Studio documentation: app overview","https:\u002F\u002Flmstudio.ai\u002Fdocs\u002Fapp",{"title":685,"url":686},"vLLM README: features, hardware and licence","https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm",{"title":688,"url":689},"vLLM documentation: GGUF support","https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Flatest\u002Ffeatures\u002Fquantization\u002Fgguf.html",{"title":691,"url":692},"GDPR Article 28: processor","https:\u002F\u002Fgdpr-info.eu\u002Fart-28-gdpr\u002F",{"title":694,"url":695},"GDPR Article 32: security of processing","https:\u002F\u002Fgdpr-info.eu\u002Fart-32-gdpr\u002F","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllama-cpp\u002Fog.jpg",[115,116,117],"llama.cpp review: the local engine under Ollama and LM Studio","llama.cpp runs open models in plain C and C++ on Metal, CUDA, Vulkan or the CPU. I cover GGUF quants, llama-server and where it falls short.","Cover art for the llama.cpp review: one GGUF file fans out to Metal, CUDA, Vulkan and CPU, then one OpenAI-style API.","Local inference runtime",{"slug":704,"published":629,"minutes":705,"category":7,"tags":706,"keywords":712,"about":720,"sources":732,"cover":802,"og":803,"expertise":113,"locales":804,"lang":115,"title":805,"description":806,"coverAlt":807,"url":744,"pricing":808,"kind":707},"opik",7,[707,708,709,710,711],"LLM observability","Tracing","Evaluation","OpenTelemetry","Self-hosting",[704,713,714,715,716,717,718,719],"opik review","opik vs langfuse","opik pricing","self-hosted llm tracing","opik opentelemetry","llm as a judge metrics","opik guardrails",[721,724,726,729],{"name":722,"url":723},"Comet ML","https:\u002F\u002Fwww.comet.com",{"name":710,"url":725},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":727,"url":728},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",{"name":730,"url":731},"Apache License","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_License",[733,736,739,742,745,748,751,754,757,760,763,766,769,772,775,778,781,784,787,790,793,796,799],{"title":734,"url":735},"Opik repository README on GitHub","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik",{"title":737,"url":738},"Opik LICENSE file: Apache License 2.0","https:\u002F\u002Fgithub.com\u002Fcomet-ml\u002Fopik\u002Fblob\u002Fmain\u002FLICENSE",{"title":740,"url":741},"Opik on PyPI: package metadata","https:\u002F\u002Fpypi.org\u002Fpypi\u002Fopik\u002Fjson",{"title":743,"url":744},"Opik product page","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fproducts\u002Fopik\u002F",{"title":746,"url":747},"Opik pricing","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fpricing\u002F",{"title":749,"url":750},"Opik docs: run locally with Docker Compose","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Flocal_deployment",{"title":752,"url":753},"Opik docs: Kubernetes deployment with Helm","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Fkubernetes",{"title":755,"url":756},"Opik docs: platform architecture","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fself-host\u002Farchitecture",{"title":758,"url":759},"Opik docs: tracing getting started","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Ftracing\u002Fgetting-started",{"title":761,"url":762},"Opik docs: OpenTelemetry integration","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fintegrations\u002Fopentelemetry",{"title":764,"url":765},"Opik docs: evaluation metrics overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fevaluation\u002Fmetrics\u002Foverview",{"title":767,"url":768},"Opik docs: online evaluation rules","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fonline-evaluation\u002Frules",{"title":770,"url":771},"Opik docs: Prompt Library overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Fprompt-library\u002Foverview",{"title":773,"url":774},"Opik docs: optimisation algorithms","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fdevelopment\u002Foptimization-runs\u002Falgorithms\u002Foverview",{"title":776,"url":777},"Opik docs: guardrails overview","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Foverview",{"title":779,"url":780},"Opik docs: guardrails server","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fguardrails\u002Fserver",{"title":782,"url":783},"Opik docs: SDK anonymizers","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fproduction\u002Fgateway-guardrails\u002Fanonymizers",{"title":785,"url":786},"Opik docs: data anonymization","https:\u002F\u002Fwww.comet.com\u002Fdocs\u002Fopik\u002Fadministration\u002Fdata_anonymization",{"title":788,"url":789},"Comet ML privacy policy","https:\u002F\u002Fwww.comet.com\u002Fsite\u002Fprivacy-policy\u002F",{"title":791,"url":792},"Comet Trust Center","https:\u002F\u002Ftrust.comet.com\u002F",{"title":794,"url":795},"Langfuse repository README (licence)","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse",{"title":797,"url":798},"Arize Phoenix repository README (licence)","https:\u002F\u002Fgithub.com\u002FArize-ai\u002Fphoenix",{"title":800,"url":801},"LangChain pricing (LangSmith plans)","https:\u002F\u002Fwww.langchain.com\u002Fpricing","\u002Fimages\u002Fblog\u002Fopik\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopik\u002Fog.jpg",[115,116,117],"Opik review: open-source tracing and evals, with a US-hosted cloud","Opik puts traces, LLM-as-a-judge metrics, prompt versions and an optimiser on one Apache 2.0 platform. Free to self-host, Pro cloud at $19 a month, US-hosted.","Cover art for the Opik review: a trace moves from the app through the backend to a judge and a score gate.","Apache 2.0 · free to self-host · Pro cloud from $19 a month",{"slug":810,"published":811,"minutes":812,"category":7,"tags":813,"keywords":816,"about":823,"sources":827,"cover":845,"og":846,"expertise":113,"locales":847,"lang":115,"title":848,"description":849,"coverAlt":850,"url":851,"pricing":852,"kind":702},"ollama","2026-09-29",11,[634,814,631,632,815],"Open models","Model serving",[810,817,818,819,820,821,822],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[824],{"name":825,"url":826},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[828,831,833,836,839,842],{"title":829,"url":830},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":832,"url":680},"Ollama on GitHub, with the MIT LICENSE file",{"title":834,"url":835},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":837,"url":838},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":840,"url":841},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":843,"url":844},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[115,116,117],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Cover art for the Ollama review: a request on port 11434 passes the scheduler and the engine and returns streamed tokens, with model loading and idle unloading noted below.","https:\u002F\u002Follama.com","MIT · free for personal use",1791636874669]