Tools/LLMOps & evals

DSPy review: compile your prompts against a metric, not by hand

DSPy compiles prompts from signatures, a metric and examples. What the optimisers cost in model calls, when they pay off, and when a hand-written prompt wins.

Type
Prompt optimisation framework
Pricing
MIT · free, you pay the model API

··8 min read

  • Prompt optimisation
  • LLM programs
  • MIPROv2
  • GEPA
  • Python
Cover art for the DSPy review: a signature and a metric feed an optimiser that compiles a saved program.

Key takeaways

  • DSPy replaces hand-written prompts with signatures, modules and optimisers that search for better prompts against a metric you write.
  • An optimisation run is paid for in model calls. On MIPROv2's light setting, a one-predictor program makes roughly 750 of them before bootstrapping.
  • Write the metric first. An optimiser can only improve what the metric measures, and that function is where most of the work sits.
  • A compiled program is a JSON file of instructions and demonstrations, so it needs version control, and its demonstrations are your training data.
  • For one prompt on a stable model, a reviewed prompt and an eval suite are simpler. DSPy earns its place in multi-call pipelines that change.

Listen to this article

0:000:00

DSPy is an open-source Python framework that treats an LLM pipeline as code. You declare inputs and outputs, pick a module, and let an optimiser search for the instructions and examples that score best on a metric you wrote. The verdict up front: use it for a multi-call pipeline you can measure and expect to change. Skip it for one prompt on a stable model – a reviewed prompt and an eval suite are simpler.

What it is

DSPy calls itself the framework for programming, not prompting, language models. Signatures declare inputs and outputs. Modules decide how the model is asked, from a plain prediction to step-by-step reasoning or a tool loop. Optimisers compile a program against a metric by changing its instructions and examples. It sits between prompt engineering and fine-tuning: the prompt becomes an artefact the optimiser writes. For the wider choice, read my decision guide on prompting, retrieval and fine-tuning.

  • MIT licence, free to use. It installs with pip install dspy and needs Python 3.10 or newer.
  • Latest release 3.4.0, published on 25 September. The repository showed 38.6k stars when I checked.
  • 14 optimisers in the API reference, from BootstrapFewShot to MIPROv2, GEPA and BootstrapFinetune.
  • No hosted service that I could find. It runs in your process and calls the provider you configure.

How it works

A module is built from a signature. At run time the adapter turns the signature into the system message, the model answers, and DSPy parses the output fields. Without an optimiser, the program is only as good as its wording. The optimiser adds a metric, a Python function that scores one prediction, usually from 0.0 to 1.0, and example inputs. It runs the program over the examples, proposes new instructions and demonstrations, keeps the best combination and returns a compiled program.

How DSPy compiles a programA signature and a module form the program. A metric and a set of examples feed the optimiser, which searches instructions and demonstrations and returns a compiled program. The compiled program can be saved as a JSON file and called in the application in place of the original module.Compile a programsame program, better promptsMetricscores one answerSignatureinputs and outputsModulePredict, CoT, ReActOptimiserMIPROv2, GEPACompiled programsaved as JSONExamplestrain and validationThe optimiser runs the program on the examples,scores each answer, and keeps the best combination.Then the compiled program replaces the module.
The metric and examples feed the optimiser, which compiles the signature and module into a program you can save.

Getting started

import os
import dspy

lm = dspy.LM("openai/gpt-5-nano", api_key=os.environ["OPENAI_API_KEY"])
dspy.configure(lm=lm)


class ExtractIntent(dspy.Signature):
    """Classify the customer's intent in one short email."""

    email: str = dspy.InputField()
    intent: str = dspy.OutputField(desc="one of: order, return, invoice, other")


extract = dspy.ChainOfThought(ExtractIntent)


def intent_metric(example, prediction, trace=None):
    return float(prediction.intent.strip().lower() == example.intent.strip().lower())


trainset = [
    dspy.Example(email="Where is my parcel 4411?", intent="order").with_inputs("email"),
    dspy.Example(email="I want my money back for the shoes.", intent="return").with_inputs("email"),
    # more labelled emails from your own inbox
]

optimizer = dspy.MIPROv2(metric=intent_metric, auto="light")
compiled = optimizer.compile(extract, trainset=trainset)
compiled.save("intent_v1.json")

The metric compares predicted and labelled intents, so this example needs labels. The compile call is where the cost sits: light makes hundreds of model calls, so run it on a small set, check the saved file, then scale up.

Signatures and modules

A signature is the contract. The string form, such as question -> answer, is shorthand. The class form adds a docstring, which becomes the instruction, and typed fields. Field order matters, because reordering inputs or outputs changes the prompt. Test every signature edit as a prompt change.

  • dspy.Predict maps inputs to outputs with a language model. Its keyword arguments go to that model.
  • dspy.ChainOfThought reasons step by step first. It adds a reasoning field you can customise.
  • dspy.ReAct runs a reason-and-act loop over tools. Its max_iters defaults to 20.

The homepage sums this up as 'same interface, different strategy'. The strategy also sets the bill: reasoning fields add output tokens to every call, and each ReAct step is another model call.

Optimisers and what they need

DSPy ships 14 optimisers, from BootstrapFewShot, which collects demonstrations, to GEPA, which rewrites instructions from the metric's feedback. All of them need a metric. The FAQ asks for a task, a metric and a few example inputs, with labels only where the metric needs them. The table covers the four I looked at most closely.

OptimiserChangesNeedsMain cost
BootstrapFewShotFew-shot demonstrationsMetric and training examplesProgram runs, one attempt per example
MIPROv2Instructions and demonstrationsMetric and training setTrials of 35 examples, plus full validation passes
GEPAInstructions, rewritten by a reflection modelMetric with feedback and a reflection modelBudget set by validation size and predictor count
BootstrapFinetuneFine-tuned model per predictorTraces and a model you can fine-tuneOne job per model, or per predictor

MIPROv2 is the usual starting point, so its arithmetic matters. Its auto setting fixes the search. For a one-predictor few-shot program, light runs about ten trials and validates on at most 100 examples, medium runs 18 trials on 300 and heavy runs 27 on 1,000. Each trial scores a 35-example minibatch, and a full validation pass runs every sixth trial, at the last trial and for the unoptimised program.

  • light: about 750 runs, 350 for trials and 400 for full passes.
  • medium: about 2,100 runs, 630 for trials and 1,500 for full passes.
  • heavy: about 7,900 runs, 945 for trials and 7,000 for full passes.

Tokens follow the calls. The FAQ, flagged as possibly out of date for DSPy 2.5 and 2.6, reports about six minutes, 3,200 calls, 2.7 million input tokens and 156,000 output tokens, for about $3 at the OpenAI pricing of the time. That is roughly 850 input and 50 output tokens per call, so the 750 calls of a light run come to about 0.6 million input tokens and 40,000 output tokens. The homepage's current example, GEPA with auto set to medium on 200 examples and gpt-5.4-mini, is listed at $2.18.

GEPA spends its budget differently. Its metric returns a score and feedback text, and a reflection model reads the examples and their scores and proposes rewritten instructions. The guide recommends a larger reflection model than the one you optimise, and the constructor requires one unless you pass a custom proposer. Light targets about six candidate prompts, and the code turns that into a metric-call budget with several full validation passes inside it.

The papers make the strongest case, each on its authors' own tasks. The DSPy paper reports pipelines beating standard few-shot prompting by over 25% and 65% for GPT-3.5 and llama2-13b-chat respectively. The MIPROv2 paper reports wins on five of seven multi-stage programs, with gains up to 13% accuracy. The GEPA paper reports 6% on average over GRPO, a reinforcement-learning baseline, with up to 35 times fewer rollouts.

Cost, deployment and data

ItemPriceWhat it covers
DSPy libraryMIT · freeInstalled with pip, runs in your process
Model calls at run timeYour provider's rateEvery call the program makes, billed per token
Vendor example run$2.18GEPA, auto medium, 200 examples, gpt-5.4-mini
Older FAQ runAbout $33,200 calls, 2.7 million input and 156,000 output tokens

Treat the compiled program as a build artefact. Save it as JSON, which the docs call safer and readable, beside the signatures in the same repository, with a version in the file name and the DSPy version pinned. It holds the signature, the demonstrations and the model for each predictor. Loading needs the same program built in code first.

Model changes trigger a recompile, because the compiler maps the program onto new prompts for the new model. Keep the old file, run the metric on both, and promote the new one only if it wins. Pin the version too: the LM page describes an auto engine that prefers a newer backend, so an upgrade can change behaviour.

The data path is the one you configure, so the processor and region questions match those for any model API. Three things also keep copies of your data. The LM cache is on by default, in memory and on disk. A saved program carries demonstrations drawn from your training data, so its JSON can hold personal data. A GEPA reflection model reads your examples and their scores, which makes it a second processor. For EU options, see my GDPR article on EU data residency.

Where it falls short

Most tasks do not need an optimiser. The FAQ concedes that for extremely simple settings a plain prompt might work just fine, and you still write the tools, retries and parsing. If you cannot write a metric that matches what a user would call correct, the optimiser improves whatever you did measure, which is not the same thing.

It beats hand-written prompts most clearly when several calls depend on each other, the model changes often, and the output can be scored. Against fine-tuning it is the cheaper and more reversible option for most teams. BootstrapFinetune compiles the program into fine-tuning jobs, but then you serve and version fine-tuned models, one per model or per predictor.

The bill and the metric are the other weak points. A light run makes hundreds of calls, heavy runs thousands, and the validation set drives most of the price. The vendor line that a small, cheap model can often match or beat a hand-prompted frontier one is a hypothesis to test on your own data.

The API is still moving. The 3.4.0 release notes list a breaking change to rlm(...) and remove the old dspy.LMRequest and dspy.LMResponse exports, so read them before each upgrade.

Verdict

Adopt DSPy when a pipeline is multi-step, measurable and changing. For one prompt on a model you never change, a reviewed prompt and a regression suite do the job. If you cannot write the metric, do not compile anything yet.

  1. Adopt it if several model calls must agree and their output can be scored automatically.
  2. Adopt it if the model or the data changes often, since recompiling beats rewriting prompts.
  3. Do not adopt it if the task is one prompt on a stable model.
  4. Do not adopt it if nobody will write and maintain the metric.

Three alternatives cover most of the rest. If you want prompts kept in code and changes gated by evals, Promptfoo is the closer fit. If the problem is explicit state and control flow, look at LangGraph, and combine the two if you need both. If the model must learn a format or a style, fine-tuning is the lever, as the decision guide explains.

Sources

Frequently asked questions

Is DSPy free to use?

The library is MIT-licensed and free. You pay for every model call, including the calls the optimiser makes while it searches. The docs' own example runs cost a few dollars.

How many examples does DSPy need?

I found no fixed minimum in the docs. The FAQ asks for a few example inputs, with labels only when the metric needs them. MIPROv2's light setting uses at most 100 examples for validation.

Can I point DSPy at a local model?

dspy.LM takes LiteLLM-style provider and model strings, so a local endpoint may work if LiteLLM can call it. I did not test one for this review, so check it on your own setup first.

Do I need to re-optimise when I change models?

Yes. The FAQ names a change of target LM as a reason to recompile, because the compiler maps the program onto new prompts. Keep the previous compiled file so you can compare and roll back.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.