# Retrieval or Fine-Tuning: How to Choose | GreenArrow

> A technical decision guide to retrieval-augmented generation versus fine-tuning: what each is good at, what it costs to run, and when a hybrid is honest.

Source: https://greenarrow.app/insights/rag-vs-fine-tuning/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Insights · Applied AI

# Retrieval or fine-tuning: how to choose

Fine-tuning is proposed far more often than it is warranted, usually because it sounds like the serious option. This is the decision as we actually make it: what each technique does, what each one costs to keep running, and the cheaper things you should exhaust first.

Applied AI Technical decision guide Reading time about 10 minutes 
 Talk to us about a build → See seven live systems 
 
 
 
 
 
 Quick answer

Use retrieval when the problem is knowledge, and fine-tuning when the problem is behaviour. Retrieval-augmented generation searches your content at question time and answers from what it finds, so it updates instantly and cites its sources. Fine-tuning adjusts model weights on examples to change style, format or a narrow classification. Most projects need retrieval, prompt engineering and structured outputs long before they need training.

## Key points

- Retrieval changes what the model is looking at. Fine-tuning changes how the model responds. Diagnose which one you have before choosing.

- Knowledge that changes, or that has to be cited and checked, belongs in retrieval. Always.

- Fine-tuning earns its place on format, voice, structured output conformance and narrow classification, not on facts.

- A fine-tune is tied to a base model checkpoint, so vendor deprecation forces you to rebuild it on somebody else's schedule.

- Prompt engineering, context engineering and schema-constrained structured outputs solve more problems than either, at a fraction of the cost.

## On this page

- What each technique actually does

- What each one is genuinely good at

- The decision table

- The cheap things to try first

- What each option costs to keep running

- How to make the decision in practice

- Hybrid architectures that work

- Frequently asked questions

Mechanics

## What each technique actually does

The two techniques are frequently compared as if they were competing products. They are not. They operate at different points in the system and they solve different classes of failure, which is why the useful question is never “which is better” but “which failure am I looking at”.

### Retrieval-augmented generation

In a retrieval system, nothing about the model changes. When a question arrives, the application searches a corpus you control, pulls back the passages most likely to contain the answer, and places them into the context window alongside the question and an instruction to answer only from what was supplied. The model then does what it is good at: reading and writing. The knowledge lives in the index, not in the weights.

In practice the search step is where the engineering effort goes. Documents have to be ingested, cleaned and split into chunks that are small enough to be precise and large enough to be coherent. Chunks are embedded into vectors for semantic search, and in almost every production system we build, that vector search runs alongside a keyword search such as BM25, because dense retrieval is poor at exact identifiers, part numbers and rare proper nouns. The combined candidate set is then reranked before anything reaches the model. Retrieval quality, not model quality, is what decides whether the answers are any good.

### Fine-tuning

Fine-tuning takes a trained model and continues training it on your own input and output pairs, adjusting the weights so that the behaviour demonstrated in those examples becomes the default. There is no search step and no citation. The change is permanent for that model and applies to every request.

Most commercial fine-tuning is now parameter-efficient: rather than updating every weight, the process trains a small set of additional parameters, commonly through low-rank adaptation, that ride on top of the frozen base model. This makes training cheaper and makes it possible to host many variants, but it does not change the fundamental property that matters here. The result is bound to a specific base checkpoint, and it encodes behaviour rather than retrievable knowledge.

### Why the distinction is not academic

A retrieval system can tell you it does not know. It can show you the paragraph it used. You can correct it by editing a document. A fine-tuned model cannot do any of those things about the knowledge it absorbed during training, because that knowledge has no address. If your organisation will ever need to answer the question “where did this come from”, that difference decides the architecture on its own.

Fit

## What each one is genuinely good at

Read this as a diagnosis aid. If your problem is not on this list, it probably belongs to prompt engineering or to the search index, and neither of those needs a training run.

### Retrieval: knowledge that moves

Prices, policies, specifications, product catalogues, contract terms, meeting records, anything with a version number. Update the document and the next answer is correct. No retraining, no release.

### Retrieval: answers that must be checkable

Regulated, clinical, legal and safety contexts where a human has to verify the source. Citation is a property of retrieval architecture, not something you can bolt on afterwards.

### Retrieval: permissioned content

Filter the corpus by the asker's entitlements before generation and the model physically cannot answer from a document that person may not read. There is no equivalent control inside a fine-tuned model.

### Fine-tuning: format and voice

A consistent house style, a fixed response shape, a domain register, terminology discipline. Behaviour you would otherwise spend hundreds of prompt tokens describing on every single call.

### Fine-tuning: narrow classification

Routing a ticket, scoring sentiment, tagging a document, detecting intent. High volume, stable definition, and a small fine-tuned model will often match a large general one at a fraction of the cost and latency.

### Fine-tuning: distillation

Teaching a small, cheap model to reproduce the behaviour of a large one on a task you have already solved expensively. This is the strongest commercial case for training, and the one clients ask for least.

Decision

## The decision table

This is the table we work through in a discovery workshop, in this order.

| What you need | Right technique | Why | What it costs to run |
|---|---|---|---|
| Answers from documents that change weekly | Retrieval | The index is updated by re-ingesting the document, and the next answer is correct immediately | Ingestion pipeline, vector store, longer prompts, periodic retrieval quality review |
| Every claim traceable to a source | Retrieval | Citation only exists if the passage was in the context window; weights have no address to cite | Same as above, plus a citation user interface and the discipline to test it |
| Answers filtered by who is asking | Retrieval | Entitlements are applied to the candidate set before generation, so unauthorised content never appears | Permission mapping into the index, and a re-sync whenever source permissions change |
| Consistent house style and tone across every reply | Fine-tuning, after prompting has been tried | Style is behaviour, and a short prompt often cannot hold enough demonstration to make it stable | Dataset curation, a training run, evaluation, and a rebuild when the base model is deprecated |
| Output that always conforms to a JSON schema | Structured outputs | Schema conformance can be enforced during decoding, so training to teach format is redundant | Schema maintenance and a validation and retry path. No training at all |
| Classifying tens of thousands of items a day | Fine-tuning a small model | A narrow task on a small model beats a general large model on cost and latency once volume is real | Labelled dataset, a held-out evaluation set, hosting, and periodic drift checks |
| A first working version this month | Prompt and context engineering | It is reversible, it costs days rather than weeks, and it produces the baseline you measure against | Almost nothing beyond the tokens, and a person who owns the prompt |
| Both correct facts and a controlled voice | Hybrid | Retrieval supplies the content, a light fine-tune or a strong system prompt supplies the manner | Both cost lines, so only take this on once each half has been justified separately |

Costs listed are the ongoing operational burden, not the build price. The build is a one-off. The burden is forever.

First resort

## The cheap things to try first

Almost every fine-tuning request that reaches us is really a request for behaviour that nobody has yet tried to specify properly. Before spending anything on training, spend a week on the two techniques that cost nearly nothing.

### Prompt engineering

A system prompt that states the role, the constraints, the refusal conditions and the output shape, followed by three or four worked examples covering the edge cases you care about. Modern models follow long, well-structured instructions far better than they did two years ago, and a great deal of the folk wisdom about prompt length is out of date. Write the prompt, run it against the evaluation set, and record the score. That number is now the bar any heavier technique has to clear.

### Context engineering

The less discussed and more valuable discipline. Context engineering is the deliberate management of everything that enters the context window: which retrieved passages, in what order, how much conversation history, which tool definitions, and what gets dropped when the window fills. Models attend unevenly across a long context, so stuffing in more material reliably makes answers worse past a certain point. Curating context is usually a bigger quality lever than changing model, and it is free.

### Structured outputs

If the complaint is “it does not always return valid JSON” or “the fields are inconsistent”, that is not a training problem. Schema-constrained decoding, exposed as structured outputs or as tool-use schemas on the major APIs, makes conformance a property of the generation process rather than a hope. Pair it with validation and a bounded retry and the class of failure disappears. This single change has removed the case for fine-tuning on several projects we were asked to quote for.

### Then measure honestly

Run the baseline. Run the candidate. Compare on the same evaluation set with the same scoring. If the improvement is within noise, the more expensive option is not the better one, it is just the more expensive one. Our approach to this on client work is described in more detail under AI consulting.

Economics

## What each option costs to keep running

The build price is the question everyone asks. The operating burden is the question that decides whether the system is still alive in eighteen months.

- Retrieval carries an ingestion pipeline. Somebody has to own connectors, re-indexing, chunking changes and the awkward document formats. This is steady, predictable work and it does not stop.

- Retrieval makes prompts longer. Retrieved passages are input tokens on every call. Prompt caching and tighter reranking bring this down, but it is a real per-request cost that scales with traffic.

- Retrieval quality decays quietly. New document types, new vocabulary and a growing corpus all degrade recall over time. Without a scheduled retrieval review, nobody notices until users stop asking.

- Fine-tuning costs are front-loaded into data. The training run is usually the cheapest line. Curating, labelling and quality-checking the dataset is where the money and the calendar go.

- A fine-tune must be rebuilt on the vendor's schedule. Base checkpoints are retired. When yours is, you repeat the dataset refresh, the training run and the full evaluation cycle, and you accept that the behaviour may shift.

- Fine-tuned knowledge goes stale silently. A retrieval system returns nothing when a document is missing. A fine-tuned model confidently answers from what it learned in March, and there is no signal that tells you it is out of date.

- Both need an evaluation set, and only one team usually builds it. The evaluation harness is the asset that lets you change anything safely. Budget it as a deliverable, not as a task somebody does at the end if there is time.

- Hosting differs more than people expect. Retrieval runs on the vendor's shared model endpoints. A fine-tune may require dedicated capacity or a higher per-token rate, which changes the unit economics at volume.

Method

## How to make the decision in practice

Seven steps, in order, and the order is the point. Each one is cheaper than the one after it.

- 01

### Write down the failure, with three real examples

Not “the answers are not good enough”. Three actual questions, the answer the system gave, and the answer it should have given. Read them together and the pattern is usually obvious: either the model did not have the information, or it had it and handled it badly. Those are different projects.

- 02

### Build the evaluation set before you build anything else

Fifty to two hundred real questions with agreed correct answers, drawn from the people who will use the system rather than from the project team. Score retrieval and generation separately, because a generation problem and a retrieval problem look identical from the outside.

- 03

### Exhaust prompt and context engineering

Rewrite the system prompt. Add worked examples. Cut the context down to what is genuinely needed. Re-run the evaluation. Record the score as your baseline. A surprising number of projects stop here, and that is a good outcome, not a disappointing one.

- 04

### Apply structured outputs to anything shaped like a format problem

Schema-constrained decoding, validation, bounded retry. If the failure was conformance, it is now gone, and you have saved a training programme.

- 05

### Add retrieval if the failure is knowledge

Ingest, chunk, index with hybrid dense and keyword search, rerank, and require citation. Test refusal explicitly: the system must say it does not know when retrieval returns nothing relevant. Re-run the evaluation and compare against the baseline.

- 06

### Only then consider fine-tuning, and scope it narrowly

Pick the smallest behaviour that still matters. Cost the dataset honestly, including the labelling. Write down what you will do when the base model is deprecated, and put a date against it. If nobody will own that rebuild, do not start.

- 07

### Re-run the whole decision in six months

Model capabilities move. A behaviour that needed training last year is often achievable in a prompt this year, and carrying a fine-tune you no longer need is a pure cost. Schedule the review.

Architecture

## Hybrid architectures that work

Once you stop treating the two as rivals, the useful combinations become obvious. Three of them come up repeatedly in production work.

### Retrieval for facts, a light fine-tune for manner

The most common hybrid. All factual content is retrieved and cited, so it stays current and checkable, while a small fine-tune handles a house style that a prompt could not hold stably. The critical rule is that no factual claim is ever expected to come from the weights. If the retrieval returns nothing, the system refuses rather than falling back on training data.

### Fine-tuned components inside a retrieval pipeline

The parts of a retrieval system that are not answer generation are excellent fine-tuning targets. Query rewriting, intent classification, routing between corpora, and reranking are all narrow, high volume, cheaply labelled from production logs, and stable enough to survive. A fine-tuned reranker frequently improves answer quality more than changing the generation model does.

### Distillation behind a retrieval front end

Where a large model has been solving a task well and the volume has grown, a small fine-tuned model trained on the large model's outputs can take over the routine traffic while the large model handles the hard cases. This is a cost engineering decision rather than a quality one, and it needs the evaluation set in place before you attempt it.

### What we see go wrong

The failure pattern is almost always the same: a team fine-tunes on a corpus of company documents in the hope of teaching the model the company, discovers that the model now writes confidently in the house style about things that are not true, and then bolts retrieval on top to fix it. The retrieval was the answer the whole time, and the fine-tune is now an expensive source of confident errors. If you take one thing from this article, take that ordering.

The same discipline applies to how these systems are governed once they are live. Model choice, training data provenance and the rebuild obligation all belong in your documentation, which is covered under AI governance, and the security implications of what goes into a retrieval corpus are covered under AI security.

Questions

## Frequently asked questions

More definitions in the glossary, and more articles in insights.

### What is the difference between RAG and fine-tuning in one sentence?

Retrieval-augmented generation searches your content at the moment a question is asked and writes the answer from the passages it finds, while fine-tuning adjusts the weights of a model on worked examples so that it behaves differently by default on every request, with no search step involved. One changes what the model is looking at. The other changes how the model responds.

### Can fine-tuning teach a model new facts?

Partly, and unreliably, which is why it is a poor instrument for knowledge. Facts absorbed during fine-tuning are spread across weights rather than stored in a retrievable form, so the model cannot cite them, cannot tell you when it is unsure of them, and will happily blend them with things it half remembers from pre-training. Worse, training on a small set of factual statements tends to make a model more confident about the shape of those statements rather than more accurate, so it invents plausible neighbours. If somebody has to be able to check where an answer came from, use retrieval.

### How much data do I need to fine-tune usefully?

For a style or format adjustment, a few hundred high-quality examples is often enough to see a real change. For a narrow classification task where you want a small model to match a large one, plan on several thousand labelled examples, plus a held-out evaluation set you did not train on. The dataset is almost always the expensive part. Most fine-tuning projects that fail do so because nobody costed the labelling, not because the training run went wrong.

### Is fine-tuning cheaper to run than retrieval?

Per request it can be, because a fine-tuned model needs a shorter prompt: the behaviour is baked in rather than described in tokens. Over the life of the system it usually is not, because you carry a dataset, a training pipeline, an evaluation set and a rebuild obligation every time the base model is deprecated. Retrieval carries an index and an ingestion pipeline instead, and those keep working when the model changes underneath them.

### What happens to my fine-tune when the vendor retires the base model?

It stops being available on the vendor's schedule, not yours. A fine-tune is an adapter on a specific base checkpoint. When that checkpoint is retired you have to run the whole training and evaluation cycle again on the successor, and the result is not guaranteed to behave the same way. That is a recurring cost and a recurring risk that belongs in the business case from the start. Retrieval systems survive a model swap with a configuration change and a re-run of the evaluation set.

### Should I try prompt engineering first?

Yes, every time. A carefully written system prompt with three or four worked examples, plus deliberate curation of what goes into the context window, resolves a large share of the problems that arrive at our door described as fine-tuning requirements. It costs a day, it is reversible, and it gives you a baseline to measure any heavier technique against. If you cannot describe the behaviour you want in a prompt, you will struggle to write a training set that teaches it.

### What are structured outputs, and do they replace fine-tuning?

Structured outputs constrain a model to emit JSON that conforms to a schema you supply, enforced during decoding rather than requested politely in a prompt. For the large category of problems that are really format conformance problems, this removes the reason to fine-tune entirely. Define the schema, validate the result, retry on failure. It is available on the main commercial APIs and on most open-weight serving stacks, and it is a great deal cheaper than a training run.

### When is a hybrid the right answer?

When you have a genuine behaviour problem and a genuine knowledge problem at the same time. The pattern we use is retrieval for the facts and a light fine-tune for the voice, the response shape or a narrow classifier inside the pipeline. Query routing, retrieval reranking and intent classification are all good fine-tuning targets because they are narrow, high volume and stable. The answer generation step is usually not.

### How do I know whether the change worked?

Build the evaluation set before you build anything else: real questions from real users, agreed correct answers, and a scoring method you can run on demand. Then measure the prompt-only baseline, and only adopt a heavier technique if it beats that baseline by a margin you would defend in a meeting. Without an evaluation set, the decision between retrieval and fine-tuning is a matter of taste, and it will be made by whoever is most senior in the room.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 04 September 2026.

Keep reading

## Related

### Twelve privacy questions before you ship AI

What a retrieval corpus contains is a privacy decision, not just an engineering one.

Read this →

### Prompt injection, explained for sign-off

Everything you put in a retrieval index becomes text your model reads and may obey.

Read this →

### AI Consulting

How we run the discovery that produces this decision on a real system.

Read this →

## Not sure which one your problem is

Send us three examples of the answers you are unhappy with. We will tell you whether it is a retrieval problem, a prompt problem or a genuine training problem, and what it would cost.

Start a conversation → 
 Read more insights
