ComparisonLLM Fine-tuning

Fine-tuning, RAG or better prompts? Diagnose the failure first

Prompting, retrieval and fine-tuning fix different problems, and choosing between them is mostly a matter of naming the failure correctly. Missing or changing knowledge points to retrieval. Inconsistent format, tone or task behavior points to fine-tuning. Weak reasoning may need a different model or a smaller task. This page gives the diagnosis, a side-by-side comparison, the combined patterns and the evaluation that proves the choice.

Reviewed 6 min read

On this page
  1. Name the failure before choosing the fix
  2. Matching each failure type to a first fix
  3. Prompting, retrieval and fine-tuning side by side
  4. Choosing by symptom
  5. Why fine-tuning is a weak way to add knowledge
  6. Combining a fine-tuned model with retrieval
  7. A hypothetical extraction task that needed both
  8. Proving the choice with a held-out comparison
  9. Questions and answers
  10. Sources

Name the failure before choosing the fix

Start with evidence, not a technique. Collect a few dozen real cases where the application gets it wrong, together with the prompt, any retrieved context and what a correct answer would have been. Then sort each case by cause. A knowledge failure means the model lacked a fact, or held an out-of-date one. A behavior failure means it knew enough but answered in the wrong format, tone or structure, or ignored a task rule. A reasoning failure means it had the facts and the format but drew the wrong conclusion over several steps.

ColdAI's fine-tuning practice starts the same way: separate knowledge from behavior, and set a prompting and retrieval baseline before deciding which behavior justifies adapting the model1. The sort usually shows a mix, and the mix decides the design.

Matching each failure type to a first fix

01Sorted failing examples02Missing facts:retrieval03Stale facts:retrieval04Format drift:fine-tune05Tone and style:prompt first06Multi-step errors:decompose07Capability gap: newmodel
  1. Sorted failing examples

    Real cases where the application was wrong, labeled by cause.

  2. Missing facts: retrieval

    The answer exists in your documents but not in the model.

  3. Stale facts: retrieval

    Policies, prices or records that change faster than any model retrains.

  4. Format drift: fine-tune

    Output structure varies even with clear instructions and examples.

  5. Tone and style: prompt first

    Try instructions and examples; fine-tune if consistency plateaus.

  6. Multi-step errors: decompose

    Split the task into checked stages or add tools before training.

  7. Capability gap: new model

    If a stronger model succeeds where yours fails, test that first.

Conceptual map from diagnosed failure types to the first fix worth testing. It is a starting point for evaluation, not a guarantee of outcome.

Prompting, retrieval and fine-tuning side by side

CriterionBetter promptingRetrieval (RAG)Fine-tuning
Fixes bestUnclear instructions, missing examples, loose output rulesKnowledge that is missing, private or changingConsistent format, tone and task behavior
What you needA test set and time to iterateClean, permissioned documents and a search indexCurated input and output examples, plus a held-out set
Time to first resultFastestModerate: ingestion, chunking and access controlSlowest: dataset curation, training and comparison
Keeping it currentEdit the promptUpdate the index; the model is unchangedRe-tune when the behavior or the base model changes
Ongoing workPrompt versioning and regression testsIndex freshness, retrieval quality and permissionsDataset versions, re-evaluation and model hosting
Main riskLong prompts that are costly and brittleWrong or unauthorized passages retrievedRegressions and memorized training data

Retrieval-augmented generation pairs a language model with a retriever over an external index, so knowledge can be updated without retraining4.

Choosing by symptom

  • If

    Answers are wrong because the information lives in your documents or systems.

    Then

    Add retrieval, and measure retrieval quality separately from answer quality.

    If the right passage is never retrieved, no amount of tuning will fix the answer.

  • If

    Facts are right but the output format, field names or structure vary between runs.

    Then

    Tighten the prompt and use structured output features first; fine-tune if consistency on held-out inputs still falls short.

    Format is behavior, and behavior is what fine-tuning changes most reliably.

  • If

    A long prompt works but is too slow or expensive at your volume.

    Then

    Fine-tune a smaller model on the outputs the long prompt produces, then compare both on the same held-out set.

    Behavior learned into the weights no longer has to be sent with every request.

  • If

    The model fails at multi-step reasoning even with the facts in front of it.

    Then

    Decompose the task, add tools or checks, or test a stronger model before training.

    Fine-tuning on worked answers transfers unevenly to problems unlike the training set.

Why fine-tuning is a weak way to add knowledge

Combining a fine-tuned model with retrieval

The choice is often both. The most common combination fine-tunes for behavior and retrieves for facts: the model learns the output schema, the house style or the task rules, while current knowledge arrives in the prompt at request time. Training examples can include retrieved passages, so the model learns to answer from supplied context and to say when the context does not contain the answer.

A second pattern uses a small fine-tuned model as a component: a classifier that routes requests, an extractor that turns retrieved documents into structured records, or a checker that validates another model's output. Each component has a narrow task, its own evaluation set and can be retrained without touching the rest.

A hypothetical extraction task that needed both

Proving the choice with a held-out comparison

  1. Freeze the evaluation set

    Combine the diagnosed failures with typical cases, drawn from sources and periods kept out of any training data.

  2. Score the best prompt

    Iterate on instructions and examples until gains stop, and record that score as the baseline.

  3. Add retrieval and rescore

    Measure answer quality and, separately, whether the right passages were retrieved.

  4. Fine-tune only for the remaining gap

    Train on examples targeting the behavior failures that survive the first two steps.

  5. Compare everything on the same set

    Include regressions, latency and running cost, and keep the simpler option if the gain does not cover the extra maintenance.

Questions and answers

Can we fine-tune a model on company documents so it knows them?

It is usually the wrong tool for that. Fine-tuning on raw documents teaches facts unreliably and fixes them at training time, so answers go stale as documents change. Index the documents for retrieval instead, with access controls, and fine-tune only if the model also needs a consistent behavior, such as a schema or a house style, that prompting cannot deliver.

Is a fine-tuned small model a substitute for a large general one?

For a narrow, well-specified task it can be: a smaller model tuned on good examples may match a larger model's output on that task at lower cost and latency. It will not match the larger model's breadth. Test it on held-out inputs from the real distribution, including unusual ones, before replacing anything.

How do we keep a fine-tuned model current?

Keep facts outside it. Put changing knowledge in retrieval so updates reach users without retraining, and reserve re-tuning for changes in the behavior itself or for moving to a new base model. Version the training data and rerun the same held-out comparison on every new release before it replaces the old one.

Can retrieval and fine-tuning be evaluated together?

Yes, and they should be. Score the full system on one frozen evaluation set, then score the parts: whether retrieval found the right passages, and whether the model used them correctly. Separating the two tells you which component to improve when the end result disappoints.

Sources

  1. LLM Fine-tuning: approach, stages and boundaries — ColdAI
  2. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs (Ovadia et al.) — arXiv · checked 10 October 2026
  3. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? (Gekhman et al.) — arXiv · checked 10 October 2026
  4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) — arXiv · checked 10 October 2026

More in LLM Fine-tuning

Back to LLM Fine-tuning

Next step

Send failing outputs and we will help sort the cause

Share examples where the application gets it wrong, with the prompt and any retrieved context. We will reply with a view on whether the gap is knowledge, behavior or reasoning, and which baseline to run before anyone trains a model.

Diagnose an LLM failure