ComparisonLLM Fine-tuning
Fine-tuning, RAG or better prompts? Diagnose the failure first
Prompting, retrieval and fine-tuning fix different problems, and choosing between them is mostly a matter of naming the failure correctly. Missing or changing knowledge points to retrieval. Inconsistent format, tone or task behavior points to fine-tuning. Weak reasoning may need a different model or a smaller task. This page gives the diagnosis, a side-by-side comparison, the combined patterns and the evaluation that proves the choice.
On this page
- Name the failure before choosing the fix
- Matching each failure type to a first fix
- Prompting, retrieval and fine-tuning side by side
- Choosing by symptom
- Why fine-tuning is a weak way to add knowledge
- Combining a fine-tuned model with retrieval
- A hypothetical extraction task that needed both
- Proving the choice with a held-out comparison
- Questions and answers
- Sources
Name the failure before choosing the fix
Start with evidence, not a technique. Collect a few dozen real cases where the application gets it wrong, together with the prompt, any retrieved context and what a correct answer would have been. Then sort each case by cause. A knowledge failure means the model lacked a fact, or held an out-of-date one. A behavior failure means it knew enough but answered in the wrong format, tone or structure, or ignored a task rule. A reasoning failure means it had the facts and the format but drew the wrong conclusion over several steps.
ColdAI's fine-tuning practice starts the same way: separate knowledge from behavior, and set a prompting and retrieval baseline before deciding which behavior justifies adapting the model1. The sort usually shows a mix, and the mix decides the design.
Matching each failure type to a first fix
- Sorted failing examples
Real cases where the application was wrong, labeled by cause.
- Missing facts: retrieval
The answer exists in your documents but not in the model.
- Stale facts: retrieval
Policies, prices or records that change faster than any model retrains.
- Format drift: fine-tune
Output structure varies even with clear instructions and examples.
- Tone and style: prompt first
Try instructions and examples; fine-tune if consistency plateaus.
- Multi-step errors: decompose
Split the task into checked stages or add tools before training.
- Capability gap: new model
If a stronger model succeeds where yours fails, test that first.
Prompting, retrieval and fine-tuning side by side
| Criterion | Better prompting | Retrieval (RAG) | Fine-tuning |
|---|---|---|---|
| Fixes best | Unclear instructions, missing examples, loose output rules | Knowledge that is missing, private or changing | Consistent format, tone and task behavior |
| What you need | A test set and time to iterate | Clean, permissioned documents and a search index | Curated input and output examples, plus a held-out set |
| Time to first result | Fastest | Moderate: ingestion, chunking and access control | Slowest: dataset curation, training and comparison |
| Keeping it current | Edit the prompt | Update the index; the model is unchanged | Re-tune when the behavior or the base model changes |
| Ongoing work | Prompt versioning and regression tests | Index freshness, retrieval quality and permissions | Dataset versions, re-evaluation and model hosting |
| Main risk | Long prompts that are costly and brittle | Wrong or unauthorized passages retrieved | Regressions and memorized training data |
Retrieval-augmented generation pairs a language model with a retriever over an external index, so knowledge can be updated without retraining4.
Choosing by symptom
- If
Answers are wrong because the information lives in your documents or systems.
ThenAdd retrieval, and measure retrieval quality separately from answer quality.
If the right passage is never retrieved, no amount of tuning will fix the answer.
- If
Facts are right but the output format, field names or structure vary between runs.
ThenTighten the prompt and use structured output features first; fine-tune if consistency on held-out inputs still falls short.
Format is behavior, and behavior is what fine-tuning changes most reliably.
- If
A long prompt works but is too slow or expensive at your volume.
ThenFine-tune a smaller model on the outputs the long prompt produces, then compare both on the same held-out set.
Behavior learned into the weights no longer has to be sent with every request.
- If
The model fails at multi-step reasoning even with the facts in front of it.
ThenDecompose the task, add tools or checks, or test a stronger model before training.
Fine-tuning on worked answers transfers unevenly to problems unlike the training set.
Why fine-tuning is a weak way to add knowledge
Combining a fine-tuned model with retrieval
The choice is often both. The most common combination fine-tunes for behavior and retrieves for facts: the model learns the output schema, the house style or the task rules, while current knowledge arrives in the prompt at request time. Training examples can include retrieved passages, so the model learns to answer from supplied context and to say when the context does not contain the answer.
A second pattern uses a small fine-tuned model as a component: a classifier that routes requests, an extractor that turns retrieved documents into structured records, or a checker that validates another model's output. Each component has a narrow task, its own evaluation set and can be retrained without touching the rest.
A hypothetical extraction task that needed both
Proving the choice with a held-out comparison
Freeze the evaluation set
Combine the diagnosed failures with typical cases, drawn from sources and periods kept out of any training data.
Score the best prompt
Iterate on instructions and examples until gains stop, and record that score as the baseline.
Add retrieval and rescore
Measure answer quality and, separately, whether the right passages were retrieved.
Fine-tune only for the remaining gap
Train on examples targeting the behavior failures that survive the first two steps.
Compare everything on the same set
Include regressions, latency and running cost, and keep the simpler option if the gain does not cover the extra maintenance.
Questions and answers
Can we fine-tune a model on company documents so it knows them?
It is usually the wrong tool for that. Fine-tuning on raw documents teaches facts unreliably and fixes them at training time, so answers go stale as documents change. Index the documents for retrieval instead, with access controls, and fine-tune only if the model also needs a consistent behavior, such as a schema or a house style, that prompting cannot deliver.
Is a fine-tuned small model a substitute for a large general one?
For a narrow, well-specified task it can be: a smaller model tuned on good examples may match a larger model's output on that task at lower cost and latency. It will not match the larger model's breadth. Test it on held-out inputs from the real distribution, including unusual ones, before replacing anything.
How do we keep a fine-tuned model current?
Keep facts outside it. Put changing knowledge in retrieval so updates reach users without retraining, and reserve re-tuning for changes in the behavior itself or for moving to a new base model. Version the training data and rerun the same held-out comparison on every new release before it replaces the old one.
Can retrieval and fine-tuning be evaluated together?
Yes, and they should be. Score the full system on one frozen evaluation set, then score the parts: whether retrieval found the right passages, and whether the model used them correctly. Separating the two tells you which component to improve when the end result disappoints.
Sources
- LLM Fine-tuning: approach, stages and boundaries — ColdAI
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs (Ovadia et al.) — arXiv · checked 10 October 2026
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? (Gekhman et al.) — arXiv · checked 10 October 2026
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) — arXiv · checked 10 October 2026