Insights · Applied AI

Retrieval or fine-tuning: how to choose

Fine-tuning is proposed far more often than it is warranted, usually because it sounds like the serious option. This is the decision as we actually make it: what each technique does, what each one costs to keep running, and the cheaper things you should exhaust first.

Applied AITechnical decision guideReading time about 10 minutes
Quick answer

Use retrieval when the problem is knowledge, and fine-tuning when the problem is behaviour. Retrieval-augmented generation searches your content at question time and answers from what it finds, so it updates instantly and cites its sources. Fine-tuning adjusts model weights on examples to change style, format or a narrow classification. Most projects need retrieval, prompt engineering and structured outputs long before they need training.

Mechanics

What each technique actually does

The two techniques are frequently compared as if they were competing products. They are not. They operate at different points in the system and they solve different classes of failure, which is why the useful question is never “which is better” but “which failure am I looking at”.

Retrieval-augmented generation

In a retrieval system, nothing about the model changes. When a question arrives, the application searches a corpus you control, pulls back the passages most likely to contain the answer, and places them into the context window alongside the question and an instruction to answer only from what was supplied. The model then does what it is good at: reading and writing. The knowledge lives in the index, not in the weights.

In practice the search step is where the engineering effort goes. Documents have to be ingested, cleaned and split into chunks that are small enough to be precise and large enough to be coherent. Chunks are embedded into vectors for semantic search, and in almost every production system we build, that vector search runs alongside a keyword search such as BM25, because dense retrieval is poor at exact identifiers, part numbers and rare proper nouns. The combined candidate set is then reranked before anything reaches the model. Retrieval quality, not model quality, is what decides whether the answers are any good.

Fine-tuning

Fine-tuning takes a trained model and continues training it on your own input and output pairs, adjusting the weights so that the behaviour demonstrated in those examples becomes the default. There is no search step and no citation. The change is permanent for that model and applies to every request.

Most commercial fine-tuning is now parameter-efficient: rather than updating every weight, the process trains a small set of additional parameters, commonly through low-rank adaptation, that ride on top of the frozen base model. This makes training cheaper and makes it possible to host many variants, but it does not change the fundamental property that matters here. The result is bound to a specific base checkpoint, and it encodes behaviour rather than retrievable knowledge.

Why the distinction is not academic

A retrieval system can tell you it does not know. It can show you the paragraph it used. You can correct it by editing a document. A fine-tuned model cannot do any of those things about the knowledge it absorbed during training, because that knowledge has no address. If your organisation will ever need to answer the question “where did this come from”, that difference decides the architecture on its own.

Fit

What each one is genuinely good at

Read this as a diagnosis aid. If your problem is not on this list, it probably belongs to prompt engineering or to the search index, and neither of those needs a training run.

Retrieval: knowledge that moves

Prices, policies, specifications, product catalogues, contract terms, meeting records, anything with a version number. Update the document and the next answer is correct. No retraining, no release.

Retrieval: answers that must be checkable

Regulated, clinical, legal and safety contexts where a human has to verify the source. Citation is a property of retrieval architecture, not something you can bolt on afterwards.

Retrieval: permissioned content

Filter the corpus by the asker's entitlements before generation and the model physically cannot answer from a document that person may not read. There is no equivalent control inside a fine-tuned model.

Fine-tuning: format and voice

A consistent house style, a fixed response shape, a domain register, terminology discipline. Behaviour you would otherwise spend hundreds of prompt tokens describing on every single call.

Fine-tuning: narrow classification

Routing a ticket, scoring sentiment, tagging a document, detecting intent. High volume, stable definition, and a small fine-tuned model will often match a large general one at a fraction of the cost and latency.

Fine-tuning: distillation

Teaching a small, cheap model to reproduce the behaviour of a large one on a task you have already solved expensively. This is the strongest commercial case for training, and the one clients ask for least.

Decision

The decision table

This is the table we work through in a discovery workshop, in this order.

What you needRight techniqueWhyWhat it costs to run
Answers from documents that change weeklyRetrievalThe index is updated by re-ingesting the document, and the next answer is correct immediatelyIngestion pipeline, vector store, longer prompts, periodic retrieval quality review
Every claim traceable to a sourceRetrievalCitation only exists if the passage was in the context window; weights have no address to citeSame as above, plus a citation user interface and the discipline to test it
Answers filtered by who is askingRetrievalEntitlements are applied to the candidate set before generation, so unauthorised content never appearsPermission mapping into the index, and a re-sync whenever source permissions change
Consistent house style and tone across every replyFine-tuning, after prompting has been triedStyle is behaviour, and a short prompt often cannot hold enough demonstration to make it stableDataset curation, a training run, evaluation, and a rebuild when the base model is deprecated
Output that always conforms to a JSON schemaStructured outputsSchema conformance can be enforced during decoding, so training to teach format is redundantSchema maintenance and a validation and retry path. No training at all
Classifying tens of thousands of items a dayFine-tuning a small modelA narrow task on a small model beats a general large model on cost and latency once volume is realLabelled dataset, a held-out evaluation set, hosting, and periodic drift checks
A first working version this monthPrompt and context engineeringIt is reversible, it costs days rather than weeks, and it produces the baseline you measure againstAlmost nothing beyond the tokens, and a person who owns the prompt
Both correct facts and a controlled voiceHybridRetrieval supplies the content, a light fine-tune or a strong system prompt supplies the mannerBoth cost lines, so only take this on once each half has been justified separately

Costs listed are the ongoing operational burden, not the build price. The build is a one-off. The burden is forever.

First resort

The cheap things to try first

Almost every fine-tuning request that reaches us is really a request for behaviour that nobody has yet tried to specify properly. Before spending anything on training, spend a week on the two techniques that cost nearly nothing.

Prompt engineering

A system prompt that states the role, the constraints, the refusal conditions and the output shape, followed by three or four worked examples covering the edge cases you care about. Modern models follow long, well-structured instructions far better than they did two years ago, and a great deal of the folk wisdom about prompt length is out of date. Write the prompt, run it against the evaluation set, and record the score. That number is now the bar any heavier technique has to clear.

Context engineering

The less discussed and more valuable discipline. Context engineering is the deliberate management of everything that enters the context window: which retrieved passages, in what order, how much conversation history, which tool definitions, and what gets dropped when the window fills. Models attend unevenly across a long context, so stuffing in more material reliably makes answers worse past a certain point. Curating context is usually a bigger quality lever than changing model, and it is free.

Structured outputs

If the complaint is “it does not always return valid JSON” or “the fields are inconsistent”, that is not a training problem. Schema-constrained decoding, exposed as structured outputs or as tool-use schemas on the major APIs, makes conformance a property of the generation process rather than a hope. Pair it with validation and a bounded retry and the class of failure disappears. This single change has removed the case for fine-tuning on several projects we were asked to quote for.

Then measure honestly

Run the baseline. Run the candidate. Compare on the same evaluation set with the same scoring. If the improvement is within noise, the more expensive option is not the better one, it is just the more expensive one. Our approach to this on client work is described in more detail under AI consulting.

Economics

What each option costs to keep running

The build price is the question everyone asks. The operating burden is the question that decides whether the system is still alive in eighteen months.

Method

How to make the decision in practice

Seven steps, in order, and the order is the point. Each one is cheaper than the one after it.

  1. 01

    Write down the failure, with three real examples

    Not “the answers are not good enough”. Three actual questions, the answer the system gave, and the answer it should have given. Read them together and the pattern is usually obvious: either the model did not have the information, or it had it and handled it badly. Those are different projects.

  2. 02

    Build the evaluation set before you build anything else

    Fifty to two hundred real questions with agreed correct answers, drawn from the people who will use the system rather than from the project team. Score retrieval and generation separately, because a generation problem and a retrieval problem look identical from the outside.

  3. 03

    Exhaust prompt and context engineering

    Rewrite the system prompt. Add worked examples. Cut the context down to what is genuinely needed. Re-run the evaluation. Record the score as your baseline. A surprising number of projects stop here, and that is a good outcome, not a disappointing one.

  4. 04

    Apply structured outputs to anything shaped like a format problem

    Schema-constrained decoding, validation, bounded retry. If the failure was conformance, it is now gone, and you have saved a training programme.

  5. 05

    Add retrieval if the failure is knowledge

    Ingest, chunk, index with hybrid dense and keyword search, rerank, and require citation. Test refusal explicitly: the system must say it does not know when retrieval returns nothing relevant. Re-run the evaluation and compare against the baseline.

  6. 06

    Only then consider fine-tuning, and scope it narrowly

    Pick the smallest behaviour that still matters. Cost the dataset honestly, including the labelling. Write down what you will do when the base model is deprecated, and put a date against it. If nobody will own that rebuild, do not start.

  7. 07

    Re-run the whole decision in six months

    Model capabilities move. A behaviour that needed training last year is often achievable in a prompt this year, and carrying a fine-tune you no longer need is a pure cost. Schedule the review.

Architecture

Hybrid architectures that work

Once you stop treating the two as rivals, the useful combinations become obvious. Three of them come up repeatedly in production work.

Retrieval for facts, a light fine-tune for manner

The most common hybrid. All factual content is retrieved and cited, so it stays current and checkable, while a small fine-tune handles a house style that a prompt could not hold stably. The critical rule is that no factual claim is ever expected to come from the weights. If the retrieval returns nothing, the system refuses rather than falling back on training data.

Fine-tuned components inside a retrieval pipeline

The parts of a retrieval system that are not answer generation are excellent fine-tuning targets. Query rewriting, intent classification, routing between corpora, and reranking are all narrow, high volume, cheaply labelled from production logs, and stable enough to survive. A fine-tuned reranker frequently improves answer quality more than changing the generation model does.

Distillation behind a retrieval front end

Where a large model has been solving a task well and the volume has grown, a small fine-tuned model trained on the large model's outputs can take over the routine traffic while the large model handles the hard cases. This is a cost engineering decision rather than a quality one, and it needs the evaluation set in place before you attempt it.

What we see go wrong

The failure pattern is almost always the same: a team fine-tunes on a corpus of company documents in the hope of teaching the model the company, discovers that the model now writes confidently in the house style about things that are not true, and then bolts retrieval on top to fix it. The retrieval was the answer the whole time, and the fine-tune is now an expensive source of confident errors. If you take one thing from this article, take that ordering.

The same discipline applies to how these systems are governed once they are live. Model choice, training data provenance and the rebuild obligation all belong in your documentation, which is covered under AI governance, and the security implications of what goes into a retrieval corpus are covered under AI security.

Questions

Frequently asked questions

More definitions in the glossary, and more articles in insights.

What is the difference between RAG and fine-tuning in one sentence?

Retrieval-augmented generation searches your content at the moment a question is asked and writes the answer from the passages it finds, while fine-tuning adjusts the weights of a model on worked examples so that it behaves differently by default on every request, with no search step involved. One changes what the model is looking at. The other changes how the model responds.

Can fine-tuning teach a model new facts?

Partly, and unreliably, which is why it is a poor instrument for knowledge. Facts absorbed during fine-tuning are spread across weights rather than stored in a retrievable form, so the model cannot cite them, cannot tell you when it is unsure of them, and will happily blend them with things it half remembers from pre-training. Worse, training on a small set of factual statements tends to make a model more confident about the shape of those statements rather than more accurate, so it invents plausible neighbours. If somebody has to be able to check where an answer came from, use retrieval.

How much data do I need to fine-tune usefully?

For a style or format adjustment, a few hundred high-quality examples is often enough to see a real change. For a narrow classification task where you want a small model to match a large one, plan on several thousand labelled examples, plus a held-out evaluation set you did not train on. The dataset is almost always the expensive part. Most fine-tuning projects that fail do so because nobody costed the labelling, not because the training run went wrong.

Is fine-tuning cheaper to run than retrieval?

Per request it can be, because a fine-tuned model needs a shorter prompt: the behaviour is baked in rather than described in tokens. Over the life of the system it usually is not, because you carry a dataset, a training pipeline, an evaluation set and a rebuild obligation every time the base model is deprecated. Retrieval carries an index and an ingestion pipeline instead, and those keep working when the model changes underneath them.

What happens to my fine-tune when the vendor retires the base model?

It stops being available on the vendor's schedule, not yours. A fine-tune is an adapter on a specific base checkpoint. When that checkpoint is retired you have to run the whole training and evaluation cycle again on the successor, and the result is not guaranteed to behave the same way. That is a recurring cost and a recurring risk that belongs in the business case from the start. Retrieval systems survive a model swap with a configuration change and a re-run of the evaluation set.

Should I try prompt engineering first?

Yes, every time. A carefully written system prompt with three or four worked examples, plus deliberate curation of what goes into the context window, resolves a large share of the problems that arrive at our door described as fine-tuning requirements. It costs a day, it is reversible, and it gives you a baseline to measure any heavier technique against. If you cannot describe the behaviour you want in a prompt, you will struggle to write a training set that teaches it.

What are structured outputs, and do they replace fine-tuning?

Structured outputs constrain a model to emit JSON that conforms to a schema you supply, enforced during decoding rather than requested politely in a prompt. For the large category of problems that are really format conformance problems, this removes the reason to fine-tune entirely. Define the schema, validate the result, retry on failure. It is available on the main commercial APIs and on most open-weight serving stacks, and it is a great deal cheaper than a training run.

When is a hybrid the right answer?

When you have a genuine behaviour problem and a genuine knowledge problem at the same time. The pattern we use is retrieval for the facts and a light fine-tune for the voice, the response shape or a narrow classifier inside the pipeline. Query routing, retrieval reranking and intent classification are all good fine-tuning targets because they are narrow, high volume and stable. The answer generation step is usually not.

How do I know whether the change worked?

Build the evaluation set before you build anything else: real questions from real users, agreed correct answers, and a scoring method you can run on demand. Then measure the prompt-only baseline, and only adopt a heavier technique if it beats that baseline by a margin you would defend in a meeting. Without an evaluation set, the decision between retrieval and fine-tuning is a matter of taste, and it will be made by whoever is most senior in the room.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Not sure which one your problem is

Send us three examples of the answers you are unhappy with. We will tell you whether it is a retrieval problem, a prompt problem or a genuine training problem, and what it would cost.