Skip to content
Happy Programming Guide
Start learning
Machine Learning

Fine-Tuning vs RAG: Which Does Your Project Need?

Fine-tuning changes how a model behaves; RAG changes what it knows at the moment of answering. The decision in one question, the trade-offs, and why many systems use both.

The inside of a computer with its parts visible

Fine-tuning and retrieval-augmented generation (RAG) are both ways to make a general language model useful for your specific situation, and they solve different problems. Fine-tuning adjusts the model’s weights so it behaves differently: a new style, a new format, a specialised skill. RAG leaves the model alone and gives it relevant documents to read at the moment it answers. The deciding question is whether your problem is about behaviour or about knowledge.

The one question#

Is the model failing because it does not know something, or because it does not do something the way you want?

  • It does not know your product documentation, last month’s policy change, the contents of your database: knowledge. Use RAG.
  • It knows enough but answers in the wrong tone, the wrong format, or lacks a skill such as writing in your domain’s conventions: behaviour. Use fine-tuning.
  • Both: use both, and do RAG first.

What RAG is#

At question time, find the few documents most relevant to the question, put them into the prompt, and ask the model to answer using them. The model’s weights never change; it simply reads what you hand it.

Python
# The shape of a RAG pipeline, with the details left to your libraries
def answer(question, index, model):
    q_vec = embed(question)                       # turn the question into a vector
    chunks = index.search(q_vec, top_k=5)         # nearest document chunks
    context = "\n\n".join(c.text for c in chunks)
    prompt = (
        "Answer the question using only the context below. "
        "If the answer is not there, say so.\n\n"
        "Context:\n" + context + "\n\nQuestion: " + question
    )
    return model.generate(prompt), [c.source for c in chunks]

Ahead of time, documents are split into chunks, each chunk is turned into a vector with an embedding model, and the vectors go into an index that supports nearest-neighbour search. That preparation is the part that takes real engineering: chunk size, overlap, which embedding model, how to keep the index in step with the documents.

Compared#

RAG Fine-tuning
Changes what the model sees how the model behaves
New information add a document, immediately live collect examples, retrain
Can cite sources yes, the retrieved chunks no
Up-front effort build the ingestion pipeline and index build a labelled dataset
Per-query cost higher: retrieval plus a longer prompt lower: short prompt, possibly a smaller model
Failure mode retrieves the wrong chunks overfits, or forgets general ability
Needs a GPU only for embeddings, and hosted options exist yes, or a hosted fine-tuning service

When RAG is clearly right#

  • Question answering over documents, especially ones that change.
  • Anything where the answer must be traceable to a source.
  • Private or per-customer data that must not be baked into a shared model.
  • Early in a project, when you do not yet know what “good” looks like well enough to label examples.

RAG also fails more gracefully: a bad answer is usually a bad retrieval, which you can inspect by looking at what was retrieved.

When fine-tuning is clearly right#

  • A consistent output format that prompting cannot make reliable enough: structured extraction, a strict schema, a house style.
  • A narrow, high-volume task where a small fine-tuned model is much cheaper per call than a large prompted one.
  • Domain language the base model handles poorly: specialised jargon, code in an internal framework.
  • Reducing prompt length: teaching the model the instructions once instead of sending them with every request.

Both#

Mature systems frequently combine them: RAG supplies the current, citable information, and a fine-tuned model reads it and responds in exactly the required shape. The order matters. Build RAG first, because it works with any model and its failures are visible. Once it is stable, the logs show whether the remaining problems are behavioural, and those logs are the beginning of a fine-tuning dataset.

The third option: a better prompt#

Before either, try a longer prompt with a few worked examples. Many “we need to fine-tune” conversations end with three examples in the system prompt doing the job. It costs nothing to try and it tells you whether the behaviour is achievable at all before you invest in a dataset.

Questions people ask#

Is RAG just a big prompt?

Mechanically, yes: retrieved text is placed in the prompt. The engineering is in choosing that text well from a corpus far too large to include, and keeping it current.

Can I fine-tune on my documents instead of building RAG?

You can, and it will help with style and vocabulary, but it will not reliably answer questions about their contents or tell you where an answer came from. RAG does both.

Does a bigger context window make RAG unnecessary?

For small corpora, you can sometimes include everything. For anything large, retrieval is still needed, and it is cheaper and faster even when the whole corpus would fit.

Which is more expensive?

RAG costs more per query and less up front; fine-tuning the reverse. At high volume on a narrow task, fine-tuning a small model usually wins on total cost.

Where to go next#

RAG explained: chunking, embeddings and retrievalRead next

Keep reading

Keep going — pick your next guide

The fastest way to improve is to read one guide, then build the thing it describes. Start with the basics, or jump straight to a project.

Ask a question or share what worked

Your email address will not be published. Required fields are marked *