Deep diveLLM Fine-tuning
LoRA, QLoRA or full fine-tuning: how each works and when to use it
LoRA, QLoRA and full fine-tuning differ in which weights change and what must fit in accelerator memory during training. LoRA trains small low-rank adapters beside frozen weights; QLoRA does the same on a base model quantized to low precision; full fine-tuning updates every weight. This page explains the mechanics, the hyperparameters worth testing first, how to limit regressions, how adapters are served and the license checks to make before training.
On this page
- Terms used when comparing fine-tuning methods
- Where memory goes during a fine-tuning run
- LoRA, QLoRA and full fine-tuning compared
- Hyperparameters to test first
- Regressions and forgetting after fine-tuning
- Serving: merge the adapter or hot-swap many
- When per-user adapters are worth evaluating
- License and data checks for an open-weight base model
- Questions and answers
- Sources
Terms used when comparing fine-tuning methods
- Low-rank adaptation (LoRA)
- Freezes the pretrained weights and learns each weight update as the product of two small matrices, so only a small fraction of parameters is trained2.
- Rank (r)
- The inner dimension of the two adapter matrices. A lower rank means smaller updates with fewer trainable parameters4.
- Alpha (lora_alpha)
- A scaling factor applied to the adapter's output; by default the update is scaled by alpha divided by rank, and the rank-stabilized variant divides by its square root instead4.
- Target modules
- The weight matrices that receive adapters, such as the attention projections or every linear layer in each transformer block.
- QLoRA
- Trains LoRA adapters while the frozen base model is stored in a 4-bit NormalFloat format, with double quantization and paged optimizers to cut memory further3.
- Merging
- Adding a trained adapter's update into the base weights to produce a standalone model with no extra inference latency4.
- Catastrophic forgetting
- Loss of abilities the base model had before fine-tuning, because training on a narrow task overwrote them.
Where memory goes during a fine-tuning run
Each method changes a different layer of this stack, which is why their hardware needs differ.
- Trainable adapters
LoRA and QLoRA train only these small matrices; full fine-tuning has no separate adapters.
- Optimizer state
Kept for every trainable parameter, so it is small for adapters and large for full fine-tuning.
- Gradients and activations
Grow with batch size and sequence length; gradient checkpointing trades compute for memory.
- Frozen base weights
Held at training precision in LoRA, quantized to low precision in QLoRA, and fully trainable in full fine-tuning.
- Accelerator memory budget
What the run must fit into, across one or several devices.
LoRA, QLoRA and full fine-tuning compared
ColdAI's fine-tuning practice works with all three methods, on Llama, Mistral and custom architectures1. The choice between them should follow evidence like this.
| Criterion | LoRA | QLoRA | Full fine-tuning |
|---|---|---|---|
| Weights updated | Adapters only; base frozen | Adapters only; base frozen and quantized | Every weight |
| Memory pressure | Moderate: full-precision base plus small optimizer state | Lowest: quantized base plus small optimizer state | Highest: optimizer state for every parameter |
| Training speed | Fast | Somewhat slower, because weights are dequantized on the fly | Slowest per step, and needs the most hardware |
| Learning capacity | Enough for most format, style and task behavior | Close to LoRA in the original paper's tests; verify on your task | Highest, for large shifts in domain or language |
| Forgetting risk | Lower | Lower | Higher; needs regression testing and often mixed-in general data |
| Artifact to store | A small adapter tied to one base version | A small adapter tied to one quantized base | A full copy of the model |
| Best fit | The default first attempt | Large base models on limited hardware | Evidence that adapters plateau below the target |
QLoRA's authors reported matching 16-bit full fine-tuning performance on their benchmarks while fine-tuning a 65B-parameter model on a single 48GB GPU3; your task needs its own comparison.
Hyperparameters to test first
Target modules often matter more than rank. The QLoRA authors found that applying adapters to all linear layers of each transformer block, not only the attention projections, was needed to match full fine-tuning3. Start there unless memory forces a narrower choice.
Rank and alpha set capacity and scale. Begin with a modest rank, keep alpha proportional to it so the effective update scale stays stable, and raise rank only if held-out results improve. With rank-stabilized scaling, higher ranks behave more predictably4.
Learning rate and epochs cause most failed runs. Adapters are usually trained with higher learning rates than full fine-tuning, but sweep a few values rather than copying one. Train for few epochs and stop when held-out loss or task scores stop improving; continuing usually memorizes the training set.
Formatting is the silent one: train with the exact chat template and special tokens the base model expects at inference, and mask the loss on prompt tokens if only the responses should be learned.
Regressions and forgetting after fine-tuning
General abilities degrade
Early signalThe tuned model handles the target task but fails simple requests the base model managed.
MitigationKeep a regression suite of general tasks. Comparisons of LoRA and full fine-tuning found LoRA forgets less of the base model's abilities but also learns less on demanding target domains5.
Safety behavior weakens
Early signalThe tuned model complies with requests the base model refused.
MitigationRe-run safety evaluations after every training run; research has shown that fine-tuning can erode safety alignment even on benign data6.
The model memorizes its training set
Early signalTraining scores keep rising while held-out scores flatten or fall, or outputs repeat training examples verbatim.
MitigationDeduplicate data, stop early, and probe for verbatim reproduction of sensitive examples.
Evaluation looks better than reality
Early signalHeld-out examples share documents, templates or customers with the training set.
MitigationSplit by source and period, as described in building a training dataset.
Serving: merge the adapter or hot-swap many
- If
You have one adapter and latency matters most.
ThenMerge it into the base weights and serve a standalone model4.
A merged model adds no inference latency, at the cost of storing a full copy.
- If
Many tasks, teams or users each need their own behavior on one base model.
ThenKeep adapters separate and load them per request on a shared base.
Serving systems have been shown to handle thousands of concurrent adapters on one base model7.
- If
The adapter was trained with QLoRA and you want to merge it.
ThenLoad the base at higher precision, merge, and re-evaluate the merged model before release.
Merging into quantized weights adds rounding error that can shift behavior.
- If
Fast rollback matters more than peak throughput.
ThenServe versioned adapters, so rollback means switching back to the previous adapter.
The base stays untouched, so the previous behavior is always one switch away.
When per-user adapters are worth evaluating
Per-user behavior is the clearest case for many adapters on one base. LinkTuned, one of ColdAI's case studies, is described as using a custom AI agent fine-tuned on each user's niche data, writing style and audience engagement patterns8. A design with that shape is where keeping small, separately versioned adapters over a shared base is worth testing against prompting with per-user examples.
License and data checks for an open-weight base model
Questions and answers
Is DPO an alternative to LoRA?
No, they answer different questions. Direct Preference Optimization is a training objective that learns from pairs of preferred and rejected responses without a separate reward model10. LoRA is a way of choosing which parameters to train. You can run DPO on LoRA adapters, typically after a supervised fine-tuning stage has taught the basic task format.
Can a LoRA adapter be moved to a different base model?
Not reliably. An adapter encodes a change relative to specific base weights, so it belongs to the exact checkpoint it was trained on. A new version of the same family, or a differently quantized copy, needs the adapter retrained and re-evaluated. Keeping the training data and configuration versioned makes that retraining routine.
Does QLoRA reduce quality compared with LoRA?
In the QLoRA paper's experiments, quality matched full-precision fine-tuning on the benchmarks reported3, but results on your task may differ. Run both methods on a small slice of your data with the same evaluation set if you have the hardware; if you do not, QLoRA is usually the practical way to tune a larger base model.
How should a fine-tuned model be evaluated after training?
Compare it with the untuned base model and the best prompting baseline on the same frozen held-out set, scored the same way. Add a regression suite for general abilities, a safety evaluation, and latency and memory measurements on the serving setup. Release only if the gain holds on unseen inputs without unacceptable regressions.
Sources
- LLM Fine-tuning: methods and model families supported — ColdAI
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al.) — arXiv · checked 10 October 2026
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al.) — arXiv · checked 10 October 2026
- PEFT documentation: LoRA conceptual guide — Hugging Face · checked 10 October 2026
- LoRA Learns Less and Forgets Less (Biderman et al.) — arXiv · checked 10 October 2026
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (Qi et al.) — arXiv · checked 10 October 2026
- S-LoRA: Serving Thousands of Concurrent LoRA Adapters (Sheng et al.) — arXiv · checked 10 October 2026
- LinkTuned: AI post automation platform — ColdAI
- Llama 3.1 Community License Agreement — Meta · checked 10 October 2026
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al.) — arXiv · checked 10 October 2026