Fine-tuning is proposed far more often than it is warranted, usually because it sounds like the serious option. This is the decision as we actually make it: what each technique does, what each one costs to keep running, and the cheaper things you should exhaust first.
Use retrieval when the problem is knowledge, and fine-tuning when the problem is behaviour. Retrieval-augmented generation searches your content at question time and answers from what it finds, so it updates instantly and cites its sources. Fine-tuning adjusts model weights on examples to change style, format or a narrow classification. Most projects need retrieval, prompt engineering and structured outputs long before they need training.
The two techniques are frequently compared as if they were competing products. They are not. They operate at different points in the system and they solve different classes of failure, which is why the useful question is never “which is better” but “which failure am I looking at”.
In a retrieval system, nothing about the model changes. When a question arrives, the application searches a corpus you control, pulls back the passages most likely to contain the answer, and places them into the context window alongside the question and an instruction to answer only from what was supplied. The model then does what it is good at: reading and writing. The knowledge lives in the index, not in the weights.
In practice the search step is where the engineering effort goes. Documents have to be ingested, cleaned and split into chunks that are small enough to be precise and large enough to be coherent. Chunks are embedded into vectors for semantic search, and in almost every production system we build, that vector search runs alongside a keyword search such as BM25, because dense retrieval is poor at exact identifiers, part numbers and rare proper nouns. The combined candidate set is then reranked before anything reaches the model. Retrieval quality, not model quality, is what decides whether the answers are any good.
Fine-tuning takes a trained model and continues training it on your own input and output pairs, adjusting the weights so that the behaviour demonstrated in those examples becomes the default. There is no search step and no citation. The change is permanent for that model and applies to every request.
Most commercial fine-tuning is now parameter-efficient: rather than updating every weight, the process trains a small set of additional parameters, commonly through low-rank adaptation, that ride on top of the frozen base model. This makes training cheaper and makes it possible to host many variants, but it does not change the fundamental property that matters here. The result is bound to a specific base checkpoint, and it encodes behaviour rather than retrievable knowledge.
A retrieval system can tell you it does not know. It can show you the paragraph it used. You can correct it by editing a document. A fine-tuned model cannot do any of those things about the knowledge it absorbed during training, because that knowledge has no address. If your organisation will ever need to answer the question “where did this come from”, that difference decides the architecture on its own.
Read this as a diagnosis aid. If your problem is not on this list, it probably belongs to prompt engineering or to the search index, and neither of those needs a training run.
Prices, policies, specifications, product catalogues, contract terms, meeting records, anything with a version number. Update the document and the next answer is correct. No retraining, no release.
Regulated, clinical, legal and safety contexts where a human has to verify the source. Citation is a property of retrieval architecture, not something you can bolt on afterwards.
Filter the corpus by the asker's entitlements before generation and the model physically cannot answer from a document that person may not read. There is no equivalent control inside a fine-tuned model.
A consistent house style, a fixed response shape, a domain register, terminology discipline. Behaviour you would otherwise spend hundreds of prompt tokens describing on every single call.
Routing a ticket, scoring sentiment, tagging a document, detecting intent. High volume, stable definition, and a small fine-tuned model will often match a large general one at a fraction of the cost and latency.
Teaching a small, cheap model to reproduce the behaviour of a large one on a task you have already solved expensively. This is the strongest commercial case for training, and the one clients ask for least.
This is the table we work through in a discovery workshop, in this order.
| What you need | Right technique | Why | What it costs to run |
|---|---|---|---|
| Answers from documents that change weekly | Retrieval | The index is updated by re-ingesting the document, and the next answer is correct immediately | Ingestion pipeline, vector store, longer prompts, periodic retrieval quality review |
| Every claim traceable to a source | Retrieval | Citation only exists if the passage was in the context window; weights have no address to cite | Same as above, plus a citation user interface and the discipline to test it |
| Answers filtered by who is asking | Retrieval | Entitlements are applied to the candidate set before generation, so unauthorised content never appears | Permission mapping into the index, and a re-sync whenever source permissions change |
| Consistent house style and tone across every reply | Fine-tuning, after prompting has been tried | Style is behaviour, and a short prompt often cannot hold enough demonstration to make it stable | Dataset curation, a training run, evaluation, and a rebuild when the base model is deprecated |
| Output that always conforms to a JSON schema | Structured outputs | Schema conformance can be enforced during decoding, so training to teach format is redundant | Schema maintenance and a validation and retry path. No training at all |
| Classifying tens of thousands of items a day | Fine-tuning a small model | A narrow task on a small model beats a general large model on cost and latency once volume is real | Labelled dataset, a held-out evaluation set, hosting, and periodic drift checks |
| A first working version this month | Prompt and context engineering | It is reversible, it costs days rather than weeks, and it produces the baseline you measure against | Almost nothing beyond the tokens, and a person who owns the prompt |
| Both correct facts and a controlled voice | Hybrid | Retrieval supplies the content, a light fine-tune or a strong system prompt supplies the manner | Both cost lines, so only take this on once each half has been justified separately |
Costs listed are the ongoing operational burden, not the build price. The build is a one-off. The burden is forever.
Almost every fine-tuning request that reaches us is really a request for behaviour that nobody has yet tried to specify properly. Before spending anything on training, spend a week on the two techniques that cost nearly nothing.
A system prompt that states the role, the constraints, the refusal conditions and the output shape, followed by three or four worked examples covering the edge cases you care about. Modern models follow long, well-structured instructions far better than they did two years ago, and a great deal of the folk wisdom about prompt length is out of date. Write the prompt, run it against the evaluation set, and record the score. That number is now the bar any heavier technique has to clear.
The less discussed and more valuable discipline. Context engineering is the deliberate management of everything that enters the context window: which retrieved passages, in what order, how much conversation history, which tool definitions, and what gets dropped when the window fills. Models attend unevenly across a long context, so stuffing in more material reliably makes answers worse past a certain point. Curating context is usually a bigger quality lever than changing model, and it is free.
If the complaint is “it does not always return valid JSON” or “the fields are inconsistent”, that is not a training problem. Schema-constrained decoding, exposed as structured outputs or as tool-use schemas on the major APIs, makes conformance a property of the generation process rather than a hope. Pair it with validation and a bounded retry and the class of failure disappears. This single change has removed the case for fine-tuning on several projects we were asked to quote for.
Run the baseline. Run the candidate. Compare on the same evaluation set with the same scoring. If the improvement is within noise, the more expensive option is not the better one, it is just the more expensive one. Our approach to this on client work is described in more detail under AI consulting.
The build price is the question everyone asks. The operating burden is the question that decides whether the system is still alive in eighteen months.
Seven steps, in order, and the order is the point. Each one is cheaper than the one after it.
Not “the answers are not good enough”. Three actual questions, the answer the system gave, and the answer it should have given. Read them together and the pattern is usually obvious: either the model did not have the information, or it had it and handled it badly. Those are different projects.
Fifty to two hundred real questions with agreed correct answers, drawn from the people who will use the system rather than from the project team. Score retrieval and generation separately, because a generation problem and a retrieval problem look identical from the outside.
Rewrite the system prompt. Add worked examples. Cut the context down to what is genuinely needed. Re-run the evaluation. Record the score as your baseline. A surprising number of projects stop here, and that is a good outcome, not a disappointing one.
Schema-constrained decoding, validation, bounded retry. If the failure was conformance, it is now gone, and you have saved a training programme.
Ingest, chunk, index with hybrid dense and keyword search, rerank, and require citation. Test refusal explicitly: the system must say it does not know when retrieval returns nothing relevant. Re-run the evaluation and compare against the baseline.
Pick the smallest behaviour that still matters. Cost the dataset honestly, including the labelling. Write down what you will do when the base model is deprecated, and put a date against it. If nobody will own that rebuild, do not start.
Model capabilities move. A behaviour that needed training last year is often achievable in a prompt this year, and carrying a fine-tune you no longer need is a pure cost. Schedule the review.
Once you stop treating the two as rivals, the useful combinations become obvious. Three of them come up repeatedly in production work.
The most common hybrid. All factual content is retrieved and cited, so it stays current and checkable, while a small fine-tune handles a house style that a prompt could not hold stably. The critical rule is that no factual claim is ever expected to come from the weights. If the retrieval returns nothing, the system refuses rather than falling back on training data.
The parts of a retrieval system that are not answer generation are excellent fine-tuning targets. Query rewriting, intent classification, routing between corpora, and reranking are all narrow, high volume, cheaply labelled from production logs, and stable enough to survive. A fine-tuned reranker frequently improves answer quality more than changing the generation model does.
Where a large model has been solving a task well and the volume has grown, a small fine-tuned model trained on the large model's outputs can take over the routine traffic while the large model handles the hard cases. This is a cost engineering decision rather than a quality one, and it needs the evaluation set in place before you attempt it.
The failure pattern is almost always the same: a team fine-tunes on a corpus of company documents in the hope of teaching the model the company, discovers that the model now writes confidently in the house style about things that are not true, and then bolts retrieval on top to fix it. The retrieval was the answer the whole time, and the fine-tune is now an expensive source of confident errors. If you take one thing from this article, take that ordering.
The same discipline applies to how these systems are governed once they are live. Model choice, training data provenance and the rebuild obligation all belong in your documentation, which is covered under AI governance, and the security implications of what goes into a retrieval corpus are covered under AI security.
More definitions in the glossary, and more articles in insights.
Retrieval-augmented generation searches your content at the moment a question is asked and writes the answer from the passages it finds, while fine-tuning adjusts the weights of a model on worked examples so that it behaves differently by default on every request, with no search step involved. One changes what the model is looking at. The other changes how the model responds.
Partly, and unreliably, which is why it is a poor instrument for knowledge. Facts absorbed during fine-tuning are spread across weights rather than stored in a retrievable form, so the model cannot cite them, cannot tell you when it is unsure of them, and will happily blend them with things it half remembers from pre-training. Worse, training on a small set of factual statements tends to make a model more confident about the shape of those statements rather than more accurate, so it invents plausible neighbours. If somebody has to be able to check where an answer came from, use retrieval.
For a style or format adjustment, a few hundred high-quality examples is often enough to see a real change. For a narrow classification task where you want a small model to match a large one, plan on several thousand labelled examples, plus a held-out evaluation set you did not train on. The dataset is almost always the expensive part. Most fine-tuning projects that fail do so because nobody costed the labelling, not because the training run went wrong.
Per request it can be, because a fine-tuned model needs a shorter prompt: the behaviour is baked in rather than described in tokens. Over the life of the system it usually is not, because you carry a dataset, a training pipeline, an evaluation set and a rebuild obligation every time the base model is deprecated. Retrieval carries an index and an ingestion pipeline instead, and those keep working when the model changes underneath them.
It stops being available on the vendor's schedule, not yours. A fine-tune is an adapter on a specific base checkpoint. When that checkpoint is retired you have to run the whole training and evaluation cycle again on the successor, and the result is not guaranteed to behave the same way. That is a recurring cost and a recurring risk that belongs in the business case from the start. Retrieval systems survive a model swap with a configuration change and a re-run of the evaluation set.
Yes, every time. A carefully written system prompt with three or four worked examples, plus deliberate curation of what goes into the context window, resolves a large share of the problems that arrive at our door described as fine-tuning requirements. It costs a day, it is reversible, and it gives you a baseline to measure any heavier technique against. If you cannot describe the behaviour you want in a prompt, you will struggle to write a training set that teaches it.
Structured outputs constrain a model to emit JSON that conforms to a schema you supply, enforced during decoding rather than requested politely in a prompt. For the large category of problems that are really format conformance problems, this removes the reason to fine-tune entirely. Define the schema, validate the result, retry on failure. It is available on the main commercial APIs and on most open-weight serving stacks, and it is a great deal cheaper than a training run.
When you have a genuine behaviour problem and a genuine knowledge problem at the same time. The pattern we use is retrieval for the facts and a light fine-tune for the voice, the response shape or a narrow classifier inside the pipeline. Query routing, retrieval reranking and intent classification are all good fine-tuning targets because they are narrow, high volume and stable. The answer generation step is usually not.
Build the evaluation set before you build anything else: real questions from real users, agreed correct answers, and a scoring method you can run on demand. Then measure the prompt-only baseline, and only adopt a heavier technique if it beats that baseline by a margin you would defend in a meeting. Without an evaluation set, the decision between retrieval and fine-tuning is a matter of taste, and it will be made by whoever is most senior in the room.
Send us three examples of the answers you are unhappy with. We will tell you whether it is a retrieval problem, a prompt problem or a genuine training problem, and what it would cost.