Skip to content
AppWizards

AI engineering · How to choose

RAG vs Fine-Tuning: When to Use Each

Start with RAG for knowledge, fine-tune for behaviour, and try prompting before either. The decision table, costs and a support assistant worked example.

Yash Rai · 24 August 2026 · updated 26 August 2026 · 8 min read

Tall library shelves packed with old books, floor to ceiling.

Use RAG when the model needs your knowledge: policies, products, prices, anything that changes. Fine-tune when the model needs a behaviour it will not learn from instructions: a strict format, a voice, a narrow skill. And before either, try a well-written prompt, because it is free and it solves more cases than the industry admits.

That is the whole decision. The rest of this article is the reasoning, the costs, and the places where the rule bends, worked through on the example we build most often: a customer support assistant.

The three options, not two

The comparison is usually framed as RAG versus fine-tuning, but every real project has three rungs, and skipping the first is the most expensive mistake on the list.

OptionWhat it changesCost shapeUndo
Prompt engineeringThe instructions the model readsHours of work, no infrastructureEdit the prompt
RAGWhat the model can look up when answeringWeeks to build, a retrieval index to runUpdate or delete documents
Fine-tuningThe model's weights themselvesData preparation, training runs, repeated per base modelRetrain

Prompt engineering is not the beginner option; it is the baseline every other option must beat. A precise prompt with good examples routinely closes half the gap teams assume needs something heavier.

RAG, retrieval augmented generation, fetches the relevant pieces of your own content at question time and hands them to the model alongside the question. The model's general ability stays as it was; its knowledge of your business becomes as fresh as your documents.

Fine-tuning continues training on your examples until the behaviour you want is in the weights. It is the only rung that changes what the model is, which is both its power and its cost.

Why RAG wins the knowledge question

A support assistant's job is answering from your return policy, your catalogue, your delivery rules. Three properties decide this in RAG's favour, and none of them is subtle.

Freshness. Your prices changed this morning. A RAG index re-embeds the changed page in minutes. A fine-tuned model knows the prices from the day its training data was cut, and retraining on every policy edit is not a workflow anyone sustains.

Citations. RAG can show which document produced the answer, which is how a person checks it and how you debug a wrong one. An answer from fine-tuned weights is an assertion with no receipt, which matters the day a customer disputes what your assistant told them.

Deletion. When a product is discontinued, you remove its page from the index and the assistant stops recommending it. Removing knowledge from weights is somewhere between hard and impossible. If your data has retention obligations, this property alone can settle the argument.

Fine-tuning a model to know your catalogue fails all three tests at once. It is the most common misuse of fine-tuning we see, and it usually arrives as "we fine-tuned it but it keeps making up products", which is exactly what the method promises.

A laptop screen showing colourful code in a dark editor.
Fine-tuning changes the model itself. That is the point, and also why it is the last resort rather than the first.Photo: Unsplash

Where fine-tuning actually earns its cost

Fine-tuning is the right tool when the target is behaviour, not facts, and the behaviour resists instructions.

  • Format compliance at scale. Output that must match a strict schema every single time, where a prompt gets you to 95 percent and the last 5 percent breaks a downstream system.
  • Voice. A brand tone that long instructions approximate but never quite hold across thousands of conversations.
  • A narrow skill on a small model. The strongest economic case in 2026: teach a small, cheap model to do one thing, classify tickets, extract fields, triage, as well as the large model does, then run it at a tenth of the cost. The large model generates training examples; the small one learns the skill. This is distillation, and it is where most of our fine-tuning work actually lands.
  • Latency and cost floors. When per-request context makes RAG too slow or too expensive at your volume, moving stable knowledge into weights can pay. This is a scale problem most businesses do not have yet.

LoRA and similar adapter methods have made the training itself cheap, hundreds of dollars rather than thousands for typical jobs on hosted platforms. The training run was never the real cost. The real cost is the dataset, hundreds to thousands of good examples that must be collected, cleaned and maintained, and the fact that the work repeats when the base model you built on is deprecated, which in this market is a when, not an if.

Where RAG projects actually fail: retrieval, not generation

"We tried RAG and the answers were bad" almost always decodes to "our retrieval returned the wrong passages, and the model faithfully answered from them". The model is the reliable half of the system. The quality levers that matter live earlier in the pipeline, and none of them is exotic.

Chunking. Documents get split into passages before indexing, and where you cut decides what can ever be found. Split a refund policy mid-clause and no query will retrieve the full rule. Chunking along a document's real structure, sections, clauses, product entries, beats fixed-size splitting in almost every corpus we have indexed.

The documents themselves. RAG is honest in an uncomfortable way: it surfaces the true state of your documentation. Contradictory policy pages, outdated price lists and answers that exist only in a senior employee's head all become visible as bad answers with citations. Cleaning the corpus is unglamorous and does more for quality than any model upgrade; it is also, conveniently, question 1 and 2 of our readiness self-test.

Retrieval tuning. Real systems combine semantic search with plain keyword matching, because customers search for order numbers and product codes, which embeddings handle poorly. A reranking step, a second, small model ordering the candidates before the answer is written, is often the single largest quality jump per rupee in the whole pipeline.

The honest fallback. The system must know when retrieval found nothing relevant and say so, rather than letting the model improvise. A measured "I do not know, here is a person" preserves trust; a fluent guess spends it.

None of this appears in vendor comparisons of RAG versus fine-tuning, yet it is where the outcome is decided. When we quote a RAG build, most of the schedule is this list, not the model integration, which is a day.

The combination, which is the usual production answer

Mature assistants often use both, each doing the job it is good at: RAG supplies the facts, and a light fine-tune holds the format and tone. The sequencing matters more than the combination: teams that fine-tune first bake in behaviour they then fight during the RAG build. Prompt first, then RAG, then fine-tune what remains, measured at every step.

Measurement is the step that makes the ladder work at all. Before touching RAG or fine-tuning, build an evaluation set, real questions with known correct answers, scored automatically. Without it you cannot tell whether the heavier method beat the lighter one, and, six months later, you cannot tell that quality has quietly dropped. We put this same discipline at the centre of how we test AI features, because it is the only reliable defence anyone has found.

What we run in our own products

Our products lean on retrieval and restraint far more than on training. Crodo, our macOS assistant, answers from what is on your screen and your connected tools, live context, which is retrieval by definition; its behaviours are held by careful prompting, not custom weights. Where we have fine-tuned, the winning case was the one described above: a small model taught one narrow skill to cut latency and cost, validated against an evaluation set the large model helped build.

That experience is why our default advice runs prompt, then RAG, then fine-tune, in that order, and why most engagements stop at the second rung.

Questions people ask

Is RAG cheaper than fine-tuning?

Usually, and more importantly its costs are steadier. RAG adds a retrieval index and some extra tokens per request, a predictable monthly line in the same range as the infrastructure numbers in our agent cost breakdown. Fine-tuning is cheap to run and expensive to feed: the dataset, the evaluation, and the retraining every time the base model changes.

When should you not use RAG?

When there is nothing to retrieve. If the task is a transformation, summarise this, classify that, rewrite this in the house style, retrieval adds cost without adding knowledge. Prompting or a light fine-tune fits transformation tasks better.

Does fine-tuning stop hallucinations?

No, and this misconception drives most wasted fine-tuning budgets. A model fine-tuned on your data still generates fluent guesses when asked something outside it. Grounding answers in retrieved documents, with an honest fallback when nothing relevant is found, is the working defence.

Is fine-tuning still relevant in 2026?

Yes, but its role has narrowed as models improved and context windows grew. The durable case is the small-model economics above, not teaching big models facts. Model-provider caching of repeated context has also absorbed part of the latency argument that used to favour weights.

Deciding for your case

Write down the ten questions your assistant must answer and where each answer lives today. If the answers live in documents that change, you are in RAG territory. If the answers never change but the output format or voice keeps drifting, you are in fine-tuning territory. If you cannot write the ten questions down yet, you are in discovery, and that is fine too; it is a two-week problem, not a two-month one.

The build itself is described on our LLM and generative AI apps page, and if the assistant needs to act rather than answer, read chatbot vs AI agent first, because that decision comes before this one.

Next step

Tell us what you want to build. We will tell you what it costs and how long it takes.

A free 30-minute call with an engineer, not a salesperson. You leave with a clear plan, a price range and an honest opinion on whether AI is the right tool for the job.