> Fine-tuning vs RAG vs prompting: what each changes (knowledge, behaviour, format), what it costs, who still offers SFT, DPO and RFT in 2026, and a decision tree.
>
> Web page: https://balazscsorba.com/blog/fine-tuning-vs-rag-vs-prompting · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/fine-tuning-vs-rag-vs-prompting.md) · [Magyar](https://balazscsorba.com/hu/blog/fine-tuning-vs-rag-vs-prompting.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: fine-tuning vs RAG, prompting vs RAG vs fine-tuning, when to fine-tune an LLM, RAG or fine-tuning, SFT vs DPO vs RFT, reinforcement fine-tuning, LLM distillation, OpenAI fine-tuning shutdown, fine-tuning decision tree, fine-tune for knowledge

[Blog](https://balazscsorba.com/blog)/LLMOps & evals

# Prompting vs RAG vs fine-tuning vs distillation: a decision guide for 2026

Fine-tuning vs RAG vs prompting: what each changes (knowledge, behaviour, format), what it costs, who still offers SFT, DPO and RFT in 2026, and a decision tree.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·12 min read

-   Fine-tuning
-   RAG
-   Prompting
-   Distillation

![Diagram: a decision flow from a tested prompt to retrieval for missing facts, supervised fine-tuning for behaviour, and distillation into a smaller model for cost.](https://balazscsorba.com/images/blog/fine-tuning-vs-rag-vs-prompting/cover.webp?v=5d127b6a8d)

## Key takeaways

-   Each lever changes something different: prompting changes instructions, RAG changes what the model can see, fine-tuning changes behaviour and format, distillation changes cost and speed.
-   Fine-tuning for knowledge is the classic mistake. Research finds RAG beats it for new facts, and examples with new knowledge can raise hallucination.
-   In 2026 the provider map has shifted: OpenAI is winding down its fine-tuning platform (new jobs end on 6 January 2027), while Google and AWS still offer managed SFT and RFT.
-   Preference tuning and RFT only pay off when you can rank outputs or write a grader, so the evaluation has to exist before the training run.
-   Start with a measured prompt, add retrieval for facts, fine-tune only a stable behaviour you can evaluate, and distill only when volume makes the cost visible.

On this page

1.  [What each lever actually changes](https://balazscsorba.com/#what-each-changes)
2.  [Who offers what in 2026](https://balazscsorba.com/#what-is-offered-in-2026)
3.  [Start with prompting, but measure it](https://balazscsorba.com/#prompting-first)
4.  [RAG for knowledge, not fine-tuning](https://balazscsorba.com/#rag-for-knowledge)
5.  [Supervised fine-tuning for behaviour and format](https://balazscsorba.com/#sft-for-behaviour)
6.  [Preference tuning and RFT: only with a signal](https://balazscsorba.com/#preference-and-rft)
7.  [Distillation: buying back cost and latency](https://balazscsorba.com/#distillation)
8.  [A decision tree](https://balazscsorba.com/#decision-tree)
9.  [Mistakes I see most often](https://balazscsorba.com/#common-mistakes)
10.  [Evaluation and maintenance checklist](https://balazscsorba.com/#evaluation-and-maintenance)
11.  [What I would do](https://balazscsorba.com/#what-i-would-do)
12.  [Sources](https://balazscsorba.com/#sources)

Every team building on LLMs reaches the same fork: the answers are not good enough, and four levers are on the table. Write a better prompt, add retrieval, fine-tune the model, or distil a smaller one. People pick by fashion. Fine-tuning feels like the serious option, so it gets chosen to fix problems it cannot fix.

The map has also moved. In May 2026 OpenAI announced it is winding down its self-serve fine-tuning platform, saying newer base models follow instructions and formats so well that prompting is now cheaper and faster. Google and AWS still offer managed tuning, and open models can be tuned anywhere. If you last compared these options a year ago, parts of your mental model are out of date.

This is the framework I use with clients. It separates the levers by what they change, compares cost, data and upkeep, lists who offers what as of 2 October 2026, names the mistakes I see most often, and ends with a decision tree and a checklist.

## What each lever actually changes

The cleanest way to think about it is to ask what is wrong with the output. There are three kinds of wrong. The model **does not know** something (knowledge). It **does the wrong thing** with what it knows (behaviour). Or it produces the right content in the **wrong shape** (format). A fourth concern, cost and latency, is separate: the output is fine but too expensive.

Lever

What it changes

Data you need

Running cost and upkeep

Typical failure

**Prompting**

Instructions, examples, output structure for one call

A handful of examples and an eval set

Longer prompts cost tokens every call; [prompt caching](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing) softens this

Prompt drift, brittle edge cases

**RAG**

What the model can see: private, fresh or large knowledge

A document corpus, plus queries to test retrieval

Index and pipeline to run; extra context tokens per call

Wrong chunk retrieved, so a confident wrong answer

**SFT**

Behaviour, tone, format, task skill

Hundreds of clean input and ideal-output pairs

Retrain when requirements or the base model change

Overfitting, narrow generalisation, stale behaviour

**Preference tuning (DPO)**

Subjective style and emphasis

Thousands of preferred and rejected pairs

Same as SFT, and the labelling is the cost

Noisy preferences train the wrong taste

**RFT**

Reasoning on tasks with a checkable answer

Dozens to hundreds of prompts, plus a grader

Grader maintenance; training billed by time

Reward hacking: the model games a weak grader

**Distillation**

Cost and latency, by moving skill into a smaller model

Teacher outputs on your real traffic

Re-distil when the task or teacher changes

Student matches the teacher on easy cases only

## Who offers what in 2026

Before choosing a method, check that your provider still sells it. This is the state I could verify from provider documentation and announcements on 2 October 2026.

**OpenAI is winding fine-tuning down**

OpenAI told developers in May 2026 that organisations that had never fine-tuned could no longer create jobs, and the community announcement gives **6 January 2027** as the date after which new jobs cannot be created at all. Inference on existing fine-tuned models continues until the underlying base model is deprecated. The documented reason: newer base models follow instructions and formats much better, prompt-based approaches are cheaper and faster, and there are fewer use cases that need fine-tuning. If you depend on an OpenAI fine-tune, plan the migration now.

Provider

Supervised (SFT)

Preference

Reinforcement (RFT)

Distillation

**OpenAI API**

GPT-4.1, 4.1-mini, 4.1-nano (winding down)

DPO on the same three models (winding down)

o4-mini only (winding down)

Not verified

**Microsoft Foundry**

GPT-4.1 family and Llama 4 Scout announced

Not verified

o4-mini, with model graders (GPT-4.1 family)

Not verified

**Google, Gemini**

Gemini 3.5 Flash, 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash and Flash-Lite

Gemini 2.5 Flash and Flash-Lite

Pre-GA preview on the same Gemini models

Via open models

**Google, open models**

Gemma, Qwen, Llama; full tuning or LoRA

Not verified

Not verified

Teacher model tunes a smaller student

**AWS Bedrock**

Amazon Nova and others; Claude 3 Haiku in us-west-2

Not verified

Nova 2 Lite, gpt-oss-20B, Qwen3 32B

Yes; teacher and student from the same family

Three practical consequences. First, "which model can I tune?" now decides more than "which method?": the frontier models most teams actually run in production are mostly not tunable, because I could find no managed fine-tuning for current Claude or GPT-5-class models. Second, RFT is real but young: Google lists it as pre-GA, and Microsoft's own RFT posts recommend starting with deterministic graders. Third, the open-weight route (LoRA on Gemma, Qwen or Llama) is the only one that is portable, and it comes with the operations burden of hosting.

## Start with prompting, but measure it

OpenAI's own optimisation guide frames the work as a loop of evals, prompt engineering and fine-tuning, and puts a baseline eval first. I agree with the order, and I would add one rule: **you may not conclude that prompting failed until you have tried the strong version of it.** That means a clear task statement, a structured output schema, three to ten real examples including hard cases, and a model that is strong enough for the task.

-   Put the format contract in a schema or structured-output mode, not in prose.
-   Use few-shot examples taken from real failures, not invented ones.
-   Cache the stable prefix. A long, repeated prompt is mostly a cost problem, and caching fixes much of it (see [cost, latency and routing](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing)).
-   Try a stronger model before tuning a weaker one. Many "fine-tuning problems" are model-size problems.
-   Write the eval set first. Without it you cannot tell if any later change helped (see [evals for product features](https://balazscsorba.com/blog/llm-evals-for-product-features)).

Google's tuning documentation gives the same split: prompting suits limited labelled data and rapid prototyping, while tuning is most effective when you have a sizeable labelled dataset (it suggests 100 examples or more) and a task that advanced prompting cannot solve. Tuning's upside it names itself: higher quality on your task and lower latency and cost thanks to shorter prompts.

## RAG for knowledge, not fine-tuning

If the model is missing facts, give it the facts at query time. This is where the most expensive mistake happens, so the evidence is worth stating. In _Fine-Tuning or Retrieval?_, Ovadia and colleagues compared unsupervised fine-tuning with RAG and found that RAG consistently outperformed it, for knowledge the model had seen in training and for entirely new knowledge; models struggled to learn new facts through fine-tuning at all. In _Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?_, Gekhman and colleagues found that examples introducing new knowledge are learned more slowly, and once learned they linearly increase the tendency to hallucinate. Their conclusion: models acquire factual knowledge mostly in pre-training, and fine-tuning teaches them to use it more efficiently.

The practical reasons are as strong as the research ones. Facts change, and a fine-tuned fact cannot be updated, cited or deleted for a GDPR request without retraining. RAG gives you sources, permissions per document and updates in minutes. The cost is a pipeline you have to build well: chunking, hybrid search and reranking, which I cover in [the RAG pipeline guide](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking) and in [RAG in 2026: hybrid, agentic or long context](https://balazscsorba.com/blog/rag-2026-hybrid-agentic-long-context).

A useful test: if the right answer exists in a document someone could hand you, it is a retrieval problem. If no document could contain it, because the problem is how to respond, it is not.

## Supervised fine-tuning for behaviour and format

SFT is the right tool when the problem is consistent behaviour. OpenAI's guidance lists classification, nuanced translation, generating content in a specific format and correcting instruction-following failures, and warns against using SFT to add entirely new knowledge. My additions from projects: strict house style, extraction into a fixed schema where prompts are too long, and compressing a 3,000-token prompt into the weights so every call gets cheaper.

**Data.** The numbers providers quote are small. OpenAI accepted as few as 10 examples and saw improvements from 50 to 100; Google's guidance says hundreds of labelled examples; AWS's Claude 3 Haiku fine-tuning took between 32 and 10,000 lines. Quantity is not the hard part. OpenAI's best-practice guide says it well: if your model has grammar, logic or style problems, check whether your data has the same problems, and a smaller amount of high-quality data beats a larger amount of low-quality data. Watch the balance too: training data with 60 percent refusals will produce a model that over-refuses.

**Method.** Parameter-efficient tuning keeps this affordable. The LoRA paper reports roughly 10,000 times fewer trainable parameters and about three times less GPU memory than full fine-tuning of GPT-3 175B, with no added inference latency. Managed services hide this; on open models you choose between LoRA and full tuning yourself.

**Maintenance.** A fine-tune is a fork of one specific base model. When that model is deprecated, you retrain, and your training data and eval set are the assets that survive. Some platforms also change how you pay: AWS requires Provisioned Throughput to run a customised Claude 3 Haiku, which turns a per-token cost into a standing one.

## Preference tuning and RFT: only with a signal

**DPO** trains on pairs: a prompt, a preferred output and a non-preferred one. OpenAI recommends it for summarising with the right emphasis and for chat with the right tone and style, and its cookbook puts the data need at thousands to tens of thousands of examples. Google recommends running SFT first and preference tuning second, and OpenAI's pipeline likewise lets SFT teach the exact wording before preference training. DPO is a polish step for taste, not a way to teach a task.

**RFT** is different in kind: the model generates answers during training and a grader scores them. OpenAI's documentation says to start small, from several dozen to a few hundred examples, to see if RFT is useful at all. It fits tasks with a checkable result such as code that passes tests, structured extraction with a known answer or constrained reasoning. Google offers string-match, Gemini-autorater and code-execution rewards, up to 16 combined; AWS reports an average 66 percent accuracy gain over base models for its RFT, which is a vendor claim from selected cases and not a guarantee. RFT does not work where the model has no initial skill or the signal is vague, and a weak grader gets exploited. Writing the grader is the real project; it is an eval by another name, so my [guide to evals](https://balazscsorba.com/blog/llm-evals-for-product-features) applies directly.

## Distillation: buying back cost and latency

Distillation is the one lever aimed at the bill. A strong teacher model generates outputs on your real inputs, and a smaller student is tuned on them. AWS describes Bedrock Model Distillation as up to five times faster and up to 75 percent cheaper than the original large model, with less than two percent accuracy loss, particularly for RAG use cases. Those are AWS's figures for its own product, so measure on your data. The constraints are instructive: teacher and student must be from the same family, and AWS itself advises using the student as is if it already performs well.

Google's documentation for open models makes the same point from the other side: distillation works best when the teacher is substantially more capable than the student, for example on multi-step reasoning, and gives smaller gains when the student is already close or the task is short-form retrieval. My rule: distil only a stable task with real volume, after you have an eval that proves the student matches the teacher. If your goal is routing between a small and a large model instead, see [typed decisions for LLM routing](https://balazscsorba.com/blog/jev-typed-decisions-llm-routing).

## A decision tree

The order of the questions is the point. Each one is cheaper to try than the next, and each answer rules out the more expensive levers.

_The questions are ordered by cost. A baseline eval comes before the first one._

Two notes on reading it. The levers combine: a production system is often a tuned or well-prompted model, with retrieval for facts, behind an eval gate. And the tree restarts whenever the base model changes, because a better model can make yesterday's fine-tune unnecessary, which is exactly what OpenAI says it has seen.

## Mistakes I see most often

-   **Fine-tuning for knowledge.** The model recites your product catalogue in training and invents prices in production. Use retrieval.
-   **Tuning before measuring.** No baseline, no eval, so nobody can say whether the tuned model is better. It is often worse on cases nobody tested.
-   **Dirty training data.** Examples copied from past model output or from inconsistent human answers teach the inconsistency.
-   **Tuning a weak model to rescue a bad task definition.** If two humans disagree on the right answer, no method will fix it.
-   **Ignoring the base-model lifecycle.** The fine-tune dies with its base model, and in 2026 the provider may stop offering tuning at all.
-   **Skipping a regression set.** Tuning for one behaviour can quietly damage others, including safety behaviour.
-   **Leaving thinking on in a tuned task.** Google advises setting the thinking budget to 0 (or the level to minimal on Gemini 3 and above) for tuned tasks, which can improve performance and cut cost.

## Evaluation and maintenance checklist

Whatever lever you pull, the same discipline applies. This is the checklist I would put in a pull request template for any change to prompts, retrieval or models.

1.  Freeze an eval set from real traffic, including failures and edge cases, before changing anything.
2.  Record the baseline for the current prompt and model, with quality, latency and cost per task.
3.  Change one lever at a time and re-run the full set, not just the cases you were looking at.
4.  Keep a regression set for behaviours you must not lose: refusals, safety, formatting.
5.  Version the training data, the grader or judge prompt and the base-model identifier next to the tuned model.
6.  Set a retraining trigger: base-model deprecation, a drift in eval score or a changed requirement.
7.  Calculate break-even: tuning and hosting cost against the tokens you save per month.

## What I would do

With a new LLM feature, I would build the eval set in week one, reach a good prompt in week two, and add retrieval the moment facts matter. I would consider fine-tuning only when a stable behaviour keeps failing after that, and only on a platform I trust to still offer it in two years, which today points to Google, AWS or open models rather than a closed API that is retreating. I would try RFT only with a deterministic grader, and distillation only when the monthly bill is large enough to pay for the project.

For most European B2B products, prompting plus retrieval plus a good eval gets 90 percent of the value. That is a judgement from my projects, not a measured number, and your eval set should overrule it. The skill is knowing which kind of wrong you are looking at.

## Sources

1.  [OpenAI community: OpenAI self-serve fine-tuning availability (wind-down announcement)](https://community.openai.com/t/openai-s-self-serve-fine-tuning-availability/1380481)
2.  [Tessl: OpenAI is shutting down self-serve fine-tuning](https://tessl.io/blog/openai-shutting-fine-tuning-signals-for-enterprise-ai)
3.  [OpenAI API docs: Model optimization](https://developers.openai.com/api/docs/guides/model-optimization)
4.  [OpenAI API docs: Supervised fine-tuning](https://developers.openai.com/api/docs/guides/supervised-fine-tuning)
5.  [OpenAI API docs: Direct preference optimization](https://developers.openai.com/api/docs/guides/direct-preference-optimization)
6.  [OpenAI API docs: Reinforcement fine-tuning](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning)
7.  [OpenAI API docs: Fine-tuning best practices](https://developers.openai.com/api/docs/guides/fine-tuning-best-practices)
8.  [OpenAI Cookbook: Choosing between SFT, DPO and RFT](https://developers.openai.com/cookbook/examples/fine_tuning_direct_preference_optimization_guide)
9.  [Google Cloud: Introduction to tuning (Gemini)](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/tune-models)
10.  [Google Cloud: About supervised fine-tuning for Gemini models](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning)
11.  [Google Cloud: About preference tuning for Gemini models](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-preference-tuning)
12.  [Google Cloud: Reward functions for reinforcement learning fine-tuning](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/reinforcement-tuning/reinforcement-tuning-job/reward-functions)
13.  [Google Cloud: Supervised and distillation fine-tuning for open models](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/open-model-tuning)
14.  [AWS: Fine-tuning for Claude 3 Haiku in Amazon Bedrock is now generally available](https://aws.amazon.com/blogs/aws/fine-tuning-for-anthropics-claude-3-haiku-model-in-amazon-bedrock-is-now-generally-available/)
15.  [AWS: Amazon Bedrock now supports reinforcement fine-tuning](https://aws.amazon.com/about-aws/whats-new/2025/12/bedrock-reinforcement-fine-tuning-66-base-models)
16.  [AWS: Amazon Bedrock Model Distillation (preview announcement)](https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/)
17.  [AWS: Bedrock reinforcement fine-tuning adds open-weight models](https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-bedrock-reinforcement-fine-tuning-openai)
18.  [Microsoft Foundry blog: What is new in Foundry fine-tuning, April 2026](https://devblogs.microsoft.com/foundry/whats-new-in-foundry-finetune-april-2026/)
19.  [Microsoft: Announcing new fine-tuning models and techniques in Azure AI Foundry](https://azure.microsoft.com/en-us/blog/announcing-new-fine-tuning-models-and-techniques-in-azure-ai-foundry/)
20.  [Ovadia et al.: Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs](https://arxiv.org/abs/2312.05934)
21.  [Gekhman et al.: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?](https://arxiv.org/abs/2405.05904)
22.  [Hu et al.: LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685)

## Frequently asked questions

Should I use RAG or fine-tuning?

Use RAG when the problem is missing or changing facts, and fine-tuning when the problem is behaviour, tone or output format. Fine-tuning is a poor way to add knowledge: a 2023 study found RAG outperformed it for both existing and new knowledge, and a 2024 study found that examples containing new facts are learned slowly and increase hallucination once learned.

When is fine-tuning worth it?

When a tested prompt with examples still fails on a stable, well-defined behaviour, you have a few hundred good examples, and you can measure the result. Typical wins are strict output formats, classification, a consistent tone and shorter prompts at high volume.

Is OpenAI fine-tuning still available in 2026?

Only for organisations that already used it. OpenAI told developers in May 2026 that it is winding the platform down: new organisations cannot create jobs, and on 6 January 2027 job creation closes for everyone. Existing fine-tuned models keep running until their base model is deprecated.

What is the difference between SFT, DPO and RFT?

Supervised fine-tuning (SFT) trains on input and ideal-output pairs. Direct preference optimization (DPO) trains on a preferred and a rejected answer per prompt. Reinforcement fine-tuning (RFT) lets the model generate answers that a grader scores, which suits tasks with a checkable result.

What is distillation and when should I use it?

Distillation uses a stronger teacher model to generate training data for a smaller student model. Use it when a task is stable and high-volume, so a cheaper model that matches the teacher on your evaluation set saves real money. AWS advises keeping the student as is if it already performs well.

How many examples do I need to fine-tune?

Providers quote small numbers to start: OpenAI accepted a minimum of 10 and saw gains from 50 to 100 examples, and Google suggests around 100 or more labelled examples. DPO usually needs far more, thousands according to OpenAI. Quality and coverage of edge cases matter more than raw count.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [Observability for LLM agents with OpenTelemetry: traces, tokens, PII and evals](https://balazscsorba.com/blog/agent-observability-opentelemetry)
-   [Claude Opus 5.5 takes #1 on Artificial Analysis, and medium effort is the real story](https://balazscsorba.com/blog/artificial-analysis-leaderboard-claude-opus-5-5)
-   [Typed decisions for LLM routing and triage: calibrated confidence with Jev](https://balazscsorba.com/blog/jev-typed-decisions-llm-routing)
-   [LLM evals for product features: from hand-read traces to a CI gate](https://balazscsorba.com/blog/llm-evals-for-product-features)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
