Blog/LLMOps & evals
Prompting vs RAG vs fine-tuning vs distillation: a decision guide for 2026
Fine-tuning vs RAG vs prompting: what each changes (knowledge, behaviour, format), what it costs, who still offers SFT, DPO and RFT in 2026, and a decision tree.
Balázs Csorba··12 min read
- Fine-tuning
- RAG
- Prompting
- Distillation

Key takeaways
- Each lever changes something different: prompting changes instructions, RAG changes what the model can see, fine-tuning changes behaviour and format, distillation changes cost and speed.
- Fine-tuning for knowledge is the classic mistake. Research finds RAG beats it for new facts, and examples with new knowledge can raise hallucination.
- In 2026 the provider map has shifted: OpenAI is winding down its fine-tuning platform (new jobs end on 6 January 2027), while Google and AWS still offer managed SFT and RFT.
- Preference tuning and RFT only pay off when you can rank outputs or write a grader, so the evaluation has to exist before the training run.
- Start with a measured prompt, add retrieval for facts, fine-tune only a stable behaviour you can evaluate, and distill only when volume makes the cost visible.
Every team building on LLMs reaches the same fork: the answers are not good enough, and four levers are on the table. Write a better prompt, add retrieval, fine-tune the model, or distil a smaller one. People pick by fashion. Fine-tuning feels like the serious option, so it gets chosen to fix problems it cannot fix.
The map has also moved. In May 2026 OpenAI announced it is winding down its self-serve fine-tuning platform, saying newer base models follow instructions and formats so well that prompting is now cheaper and faster. Google and AWS still offer managed tuning, and open models can be tuned anywhere. If you last compared these options a year ago, parts of your mental model are out of date.
This is the framework I use with clients. It separates the levers by what they change, compares cost, data and upkeep, lists who offers what as of 2 October 2026, names the mistakes I see most often, and ends with a decision tree and a checklist.
What each lever actually changes
The cleanest way to think about it is to ask what is wrong with the output. There are three kinds of wrong. The model does not know something (knowledge). It does the wrong thing with what it knows (behaviour). Or it produces the right content in the wrong shape (format). A fourth concern, cost and latency, is separate: the output is fine but too expensive.
| Lever | What it changes | Data you need | Running cost and upkeep | Typical failure |
|---|---|---|---|---|
| Prompting | Instructions, examples, output structure for one call | A handful of examples and an eval set | Longer prompts cost tokens every call; prompt caching softens this | Prompt drift, brittle edge cases |
| RAG | What the model can see: private, fresh or large knowledge | A document corpus, plus queries to test retrieval | Index and pipeline to run; extra context tokens per call | Wrong chunk retrieved, so a confident wrong answer |
| SFT | Behaviour, tone, format, task skill | Hundreds of clean input and ideal-output pairs | Retrain when requirements or the base model change | Overfitting, narrow generalisation, stale behaviour |
| Preference tuning (DPO) | Subjective style and emphasis | Thousands of preferred and rejected pairs | Same as SFT, and the labelling is the cost | Noisy preferences train the wrong taste |
| RFT | Reasoning on tasks with a checkable answer | Dozens to hundreds of prompts, plus a grader | Grader maintenance; training billed by time | Reward hacking: the model games a weak grader |
| Distillation | Cost and latency, by moving skill into a smaller model | Teacher outputs on your real traffic | Re-distil when the task or teacher changes | Student matches the teacher on easy cases only |
Who offers what in 2026
Before choosing a method, check that your provider still sells it. This is the state I could verify from provider documentation and announcements on 2 October 2026.
| Provider | Supervised (SFT) | Preference | Reinforcement (RFT) | Distillation |
|---|---|---|---|---|
| OpenAI API | GPT-4.1, 4.1-mini, 4.1-nano (winding down) | DPO on the same three models (winding down) | o4-mini only (winding down) | Not verified |
| Microsoft Foundry | GPT-4.1 family and Llama 4 Scout announced | Not verified | o4-mini, with model graders (GPT-4.1 family) | Not verified |
| Google, Gemini | Gemini 3.5 Flash, 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash and Flash-Lite | Gemini 2.5 Flash and Flash-Lite | Pre-GA preview on the same Gemini models | Via open models |
| Google, open models | Gemma, Qwen, Llama; full tuning or LoRA | Not verified | Not verified | Teacher model tunes a smaller student |
| AWS Bedrock | Amazon Nova and others; Claude 3 Haiku in us-west-2 | Not verified | Nova 2 Lite, gpt-oss-20B, Qwen3 32B | Yes; teacher and student from the same family |
Three practical consequences. First, "which model can I tune?" now decides more than "which method?": the frontier models most teams actually run in production are mostly not tunable, because I could find no managed fine-tuning for current Claude or GPT-5-class models. Second, RFT is real but young: Google lists it as pre-GA, and Microsoft's own RFT posts recommend starting with deterministic graders. Third, the open-weight route (LoRA on Gemma, Qwen or Llama) is the only one that is portable, and it comes with the operations burden of hosting.
Start with prompting, but measure it
OpenAI's own optimisation guide frames the work as a loop of evals, prompt engineering and fine-tuning, and puts a baseline eval first. I agree with the order, and I would add one rule: you may not conclude that prompting failed until you have tried the strong version of it. That means a clear task statement, a structured output schema, three to ten real examples including hard cases, and a model that is strong enough for the task.
- Put the format contract in a schema or structured-output mode, not in prose.
- Use few-shot examples taken from real failures, not invented ones.
- Cache the stable prefix. A long, repeated prompt is mostly a cost problem, and caching fixes much of it (see cost, latency and routing).
- Try a stronger model before tuning a weaker one. Many "fine-tuning problems" are model-size problems.
- Write the eval set first. Without it you cannot tell if any later change helped (see evals for product features).
Google's tuning documentation gives the same split: prompting suits limited labelled data and rapid prototyping, while tuning is most effective when you have a sizeable labelled dataset (it suggests 100 examples or more) and a task that advanced prompting cannot solve. Tuning's upside it names itself: higher quality on your task and lower latency and cost thanks to shorter prompts.
RAG for knowledge, not fine-tuning
If the model is missing facts, give it the facts at query time. This is where the most expensive mistake happens, so the evidence is worth stating. In Fine-Tuning or Retrieval?, Ovadia and colleagues compared unsupervised fine-tuning with RAG and found that RAG consistently outperformed it, for knowledge the model had seen in training and for entirely new knowledge; models struggled to learn new facts through fine-tuning at all. In Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, Gekhman and colleagues found that examples introducing new knowledge are learned more slowly, and once learned they linearly increase the tendency to hallucinate. Their conclusion: models acquire factual knowledge mostly in pre-training, and fine-tuning teaches them to use it more efficiently.
The practical reasons are as strong as the research ones. Facts change, and a fine-tuned fact cannot be updated, cited or deleted for a GDPR request without retraining. RAG gives you sources, permissions per document and updates in minutes. The cost is a pipeline you have to build well: chunking, hybrid search and reranking, which I cover in the RAG pipeline guide and in RAG in 2026: hybrid, agentic or long context.
A useful test: if the right answer exists in a document someone could hand you, it is a retrieval problem. If no document could contain it, because the problem is how to respond, it is not.
Supervised fine-tuning for behaviour and format
SFT is the right tool when the problem is consistent behaviour. OpenAI's guidance lists classification, nuanced translation, generating content in a specific format and correcting instruction-following failures, and warns against using SFT to add entirely new knowledge. My additions from projects: strict house style, extraction into a fixed schema where prompts are too long, and compressing a 3,000-token prompt into the weights so every call gets cheaper.
Data. The numbers providers quote are small. OpenAI accepted as few as 10 examples and saw improvements from 50 to 100; Google's guidance says hundreds of labelled examples; AWS's Claude 3 Haiku fine-tuning took between 32 and 10,000 lines. Quantity is not the hard part. OpenAI's best-practice guide says it well: if your model has grammar, logic or style problems, check whether your data has the same problems, and a smaller amount of high-quality data beats a larger amount of low-quality data. Watch the balance too: training data with 60 percent refusals will produce a model that over-refuses.
Method. Parameter-efficient tuning keeps this affordable. The LoRA paper reports roughly 10,000 times fewer trainable parameters and about three times less GPU memory than full fine-tuning of GPT-3 175B, with no added inference latency. Managed services hide this; on open models you choose between LoRA and full tuning yourself.
Maintenance. A fine-tune is a fork of one specific base model. When that model is deprecated, you retrain, and your training data and eval set are the assets that survive. Some platforms also change how you pay: AWS requires Provisioned Throughput to run a customised Claude 3 Haiku, which turns a per-token cost into a standing one.
Preference tuning and RFT: only with a signal
DPO trains on pairs: a prompt, a preferred output and a non-preferred one. OpenAI recommends it for summarising with the right emphasis and for chat with the right tone and style, and its cookbook puts the data need at thousands to tens of thousands of examples. Google recommends running SFT first and preference tuning second, and OpenAI's pipeline likewise lets SFT teach the exact wording before preference training. DPO is a polish step for taste, not a way to teach a task.
RFT is different in kind: the model generates answers during training and a grader scores them. OpenAI's documentation says to start small, from several dozen to a few hundred examples, to see if RFT is useful at all. It fits tasks with a checkable result such as code that passes tests, structured extraction with a known answer or constrained reasoning. Google offers string-match, Gemini-autorater and code-execution rewards, up to 16 combined; AWS reports an average 66 percent accuracy gain over base models for its RFT, which is a vendor claim from selected cases and not a guarantee. RFT does not work where the model has no initial skill or the signal is vague, and a weak grader gets exploited. Writing the grader is the real project; it is an eval by another name, so my guide to evals applies directly.
Distillation: buying back cost and latency
Distillation is the one lever aimed at the bill. A strong teacher model generates outputs on your real inputs, and a smaller student is tuned on them. AWS describes Bedrock Model Distillation as up to five times faster and up to 75 percent cheaper than the original large model, with less than two percent accuracy loss, particularly for RAG use cases. Those are AWS's figures for its own product, so measure on your data. The constraints are instructive: teacher and student must be from the same family, and AWS itself advises using the student as is if it already performs well.
Google's documentation for open models makes the same point from the other side: distillation works best when the teacher is substantially more capable than the student, for example on multi-step reasoning, and gives smaller gains when the student is already close or the task is short-form retrieval. My rule: distil only a stable task with real volume, after you have an eval that proves the student matches the teacher. If your goal is routing between a small and a large model instead, see typed decisions for LLM routing.
A decision tree
The order of the questions is the point. Each one is cheaper to try than the next, and each answer rules out the more expensive levers.
Two notes on reading it. The levers combine: a production system is often a tuned or well-prompted model, with retrieval for facts, behind an eval gate. And the tree restarts whenever the base model changes, because a better model can make yesterday's fine-tune unnecessary, which is exactly what OpenAI says it has seen.
Mistakes I see most often
- Fine-tuning for knowledge. The model recites your product catalogue in training and invents prices in production. Use retrieval.
- Tuning before measuring. No baseline, no eval, so nobody can say whether the tuned model is better. It is often worse on cases nobody tested.
- Dirty training data. Examples copied from past model output or from inconsistent human answers teach the inconsistency.
- Tuning a weak model to rescue a bad task definition. If two humans disagree on the right answer, no method will fix it.
- Ignoring the base-model lifecycle. The fine-tune dies with its base model, and in 2026 the provider may stop offering tuning at all.
- Skipping a regression set. Tuning for one behaviour can quietly damage others, including safety behaviour.
- Leaving thinking on in a tuned task. Google advises setting the thinking budget to 0 (or the level to minimal on Gemini 3 and above) for tuned tasks, which can improve performance and cut cost.
Evaluation and maintenance checklist
Whatever lever you pull, the same discipline applies. This is the checklist I would put in a pull request template for any change to prompts, retrieval or models.
- Freeze an eval set from real traffic, including failures and edge cases, before changing anything.
- Record the baseline for the current prompt and model, with quality, latency and cost per task.
- Change one lever at a time and re-run the full set, not just the cases you were looking at.
- Keep a regression set for behaviours you must not lose: refusals, safety, formatting.
- Version the training data, the grader or judge prompt and the base-model identifier next to the tuned model.
- Set a retraining trigger: base-model deprecation, a drift in eval score or a changed requirement.
- Calculate break-even: tuning and hosting cost against the tokens you save per month.
What I would do
With a new LLM feature, I would build the eval set in week one, reach a good prompt in week two, and add retrieval the moment facts matter. I would consider fine-tuning only when a stable behaviour keeps failing after that, and only on a platform I trust to still offer it in two years, which today points to Google, AWS or open models rather than a closed API that is retreating. I would try RFT only with a deterministic grader, and distillation only when the monthly bill is large enough to pay for the project.
For most European B2B products, prompting plus retrieval plus a good eval gets 90 percent of the value. That is a judgement from my projects, not a measured number, and your eval set should overrule it. The skill is knowing which kind of wrong you are looking at.
Sources
- OpenAI community: OpenAI self-serve fine-tuning availability (wind-down announcement)
- Tessl: OpenAI is shutting down self-serve fine-tuning
- OpenAI API docs: Model optimization
- OpenAI API docs: Supervised fine-tuning
- OpenAI API docs: Direct preference optimization
- OpenAI API docs: Reinforcement fine-tuning
- OpenAI API docs: Fine-tuning best practices
- OpenAI Cookbook: Choosing between SFT, DPO and RFT
- Google Cloud: Introduction to tuning (Gemini)
- Google Cloud: About supervised fine-tuning for Gemini models
- Google Cloud: About preference tuning for Gemini models
- Google Cloud: Reward functions for reinforcement learning fine-tuning
- Google Cloud: Supervised and distillation fine-tuning for open models
- AWS: Fine-tuning for Claude 3 Haiku in Amazon Bedrock is now generally available
- AWS: Amazon Bedrock now supports reinforcement fine-tuning
- AWS: Amazon Bedrock Model Distillation (preview announcement)
- AWS: Bedrock reinforcement fine-tuning adds open-weight models
- Microsoft Foundry blog: What is new in Foundry fine-tuning, April 2026
- Microsoft: Announcing new fine-tuning models and techniques in Azure AI Foundry
- Ovadia et al.: Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
- Gekhman et al.: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Hu et al.: LoRA: Low-Rank Adaptation of Large Language Models
Frequently asked questions
Should I use RAG or fine-tuning?
Use RAG when the problem is missing or changing facts, and fine-tuning when the problem is behaviour, tone or output format. Fine-tuning is a poor way to add knowledge: a 2023 study found RAG outperformed it for both existing and new knowledge, and a 2024 study found that examples containing new facts are learned slowly and increase hallucination once learned.
When is fine-tuning worth it?
When a tested prompt with examples still fails on a stable, well-defined behaviour, you have a few hundred good examples, and you can measure the result. Typical wins are strict output formats, classification, a consistent tone and shorter prompts at high volume.
Is OpenAI fine-tuning still available in 2026?
Only for organisations that already used it. OpenAI told developers in May 2026 that it is winding the platform down: new organisations cannot create jobs, and on 6 January 2027 job creation closes for everyone. Existing fine-tuned models keep running until their base model is deprecated.
What is the difference between SFT, DPO and RFT?
Supervised fine-tuning (SFT) trains on input and ideal-output pairs. Direct preference optimization (DPO) trains on a preferred and a rejected answer per prompt. Reinforcement fine-tuning (RFT) lets the model generate answers that a grader scores, which suits tasks with a checkable result.
What is distillation and when should I use it?
Distillation uses a stronger teacher model to generate training data for a smaller student model. Use it when a task is stable and high-volume, so a cheaper model that matches the teacher on your evaluation set saves real money. AWS advises keeping the student as is if it already performs well.
How many examples do I need to fine-tune?
Providers quote small numbers to start: OpenAI accepted a minimum of 10 and saw gains from 50 to 100 examples, and Google suggests around 100 or more labelled examples. DPO usually needs far more, thousands according to OpenAI. Quality and coverage of edge cases matter more than raw count.