Fine-tuning and retrieval-augmented generation (RAG) are both ways to make a general language model useful for your specific situation, and they solve different problems. Fine-tuning adjusts the model’s weights so it behaves differently: a new style, a new format, a specialised skill. RAG leaves the model alone and gives it relevant documents to read at the moment it answers. The deciding question is whether your problem is about behaviour or about knowledge.
The one question#
Is the model failing because it does not know something, or because it does not do something the way you want?
- It does not know your product documentation, last month’s policy change, the contents of your database: knowledge. Use RAG.
- It knows enough but answers in the wrong tone, the wrong format, or lacks a skill such as writing in your domain’s conventions: behaviour. Use fine-tuning.
- Both: use both, and do RAG first.
What RAG is#
At question time, find the few documents most relevant to the question, put them into the prompt, and ask the model to answer using them. The model’s weights never change; it simply reads what you hand it.
# The shape of a RAG pipeline, with the details left to your libraries
def answer(question, index, model):
q_vec = embed(question) # turn the question into a vector
chunks = index.search(q_vec, top_k=5) # nearest document chunks
context = "\n\n".join(c.text for c in chunks)
prompt = (
"Answer the question using only the context below. "
"If the answer is not there, say so.\n\n"
"Context:\n" + context + "\n\nQuestion: " + question
)
return model.generate(prompt), [c.source for c in chunks]
Ahead of time, documents are split into chunks, each chunk is turned into a vector with an embedding model, and the vectors go into an index that supports nearest-neighbour search. That preparation is the part that takes real engineering: chunk size, overlap, which embedding model, how to keep the index in step with the documents.
Compared#
| RAG | Fine-tuning | |
|---|---|---|
| Changes | what the model sees | how the model behaves |
| New information | add a document, immediately live | collect examples, retrain |
| Can cite sources | yes, the retrieved chunks | no |
| Up-front effort | build the ingestion pipeline and index | build a labelled dataset |
| Per-query cost | higher: retrieval plus a longer prompt | lower: short prompt, possibly a smaller model |
| Failure mode | retrieves the wrong chunks | overfits, or forgets general ability |
| Needs a GPU | only for embeddings, and hosted options exist | yes, or a hosted fine-tuning service |
When RAG is clearly right#
- Question answering over documents, especially ones that change.
- Anything where the answer must be traceable to a source.
- Private or per-customer data that must not be baked into a shared model.
- Early in a project, when you do not yet know what “good” looks like well enough to label examples.
RAG also fails more gracefully: a bad answer is usually a bad retrieval, which you can inspect by looking at what was retrieved.
When fine-tuning is clearly right#
- A consistent output format that prompting cannot make reliable enough: structured extraction, a strict schema, a house style.
- A narrow, high-volume task where a small fine-tuned model is much cheaper per call than a large prompted one.
- Domain language the base model handles poorly: specialised jargon, code in an internal framework.
- Reducing prompt length: teaching the model the instructions once instead of sending them with every request.
Both#
Mature systems frequently combine them: RAG supplies the current, citable information, and a fine-tuned model reads it and responds in exactly the required shape. The order matters. Build RAG first, because it works with any model and its failures are visible. Once it is stable, the logs show whether the remaining problems are behavioural, and those logs are the beginning of a fine-tuning dataset.
The third option: a better prompt#
Before either, try a longer prompt with a few worked examples. Many “we need to fine-tune” conversations end with three examples in the system prompt doing the job. It costs nothing to try and it tells you whether the behaviour is achievable at all before you invest in a dataset.
Questions people ask#
Is RAG just a big prompt?
Mechanically, yes: retrieved text is placed in the prompt. The engineering is in choosing that text well from a corpus far too large to include, and keeping it current.
Can I fine-tune on my documents instead of building RAG?
You can, and it will help with style and vocabulary, but it will not reliably answer questions about their contents or tell you where an answer came from. RAG does both.
Does a bigger context window make RAG unnecessary?
For small corpora, you can sometimes include everything. For anything large, retrieval is still needed, and it is cheaper and faster even when the whole corpus would fit.
Which is more expensive?
RAG costs more per query and less up front; fine-tuning the reverse. At high volume on a narrow task, fine-tuning a small model usually wins on total cost.
Where to go next#
- RAG explained — the pipeline in detail.
- What is fine-tuning? — with a working example.
- Large language models explained — what both techniques are adjusting.