Deep diveLLM Fine-tuning

LoRA, QLoRA or full fine-tuning: how each works and when to use it

LoRA, QLoRA and full fine-tuning differ in which weights change and what must fit in accelerator memory during training. LoRA trains small low-rank adapters beside frozen weights; QLoRA does the same on a base model quantized to low precision; full fine-tuning updates every weight. This page explains the mechanics, the hyperparameters worth testing first, how to limit regressions, how adapters are served and the license checks to make before training.

Reviewed 7 min read

On this page
  1. Terms used when comparing fine-tuning methods
  2. Where memory goes during a fine-tuning run
  3. LoRA, QLoRA and full fine-tuning compared
  4. Hyperparameters to test first
  5. Regressions and forgetting after fine-tuning
  6. Serving: merge the adapter or hot-swap many
  7. When per-user adapters are worth evaluating
  8. License and data checks for an open-weight base model
  9. Questions and answers
  10. Sources

Terms used when comparing fine-tuning methods

Low-rank adaptation (LoRA)
Freezes the pretrained weights and learns each weight update as the product of two small matrices, so only a small fraction of parameters is trained2.
Rank (r)
The inner dimension of the two adapter matrices. A lower rank means smaller updates with fewer trainable parameters4.
Alpha (lora_alpha)
A scaling factor applied to the adapter's output; by default the update is scaled by alpha divided by rank, and the rank-stabilized variant divides by its square root instead4.
Target modules
The weight matrices that receive adapters, such as the attention projections or every linear layer in each transformer block.
QLoRA
Trains LoRA adapters while the frozen base model is stored in a 4-bit NormalFloat format, with double quantization and paged optimizers to cut memory further3.
Merging
Adding a trained adapter's update into the base weights to produce a standalone model with no extra inference latency4.
Catastrophic forgetting
Loss of abilities the base model had before fine-tuning, because training on a narrow task overwrote them.

Where memory goes during a fine-tuning run

Each method changes a different layer of this stack, which is why their hardware needs differ.

Trainable adapters01Optimizer state02Gradients and activations03Frozen base weights04Accelerator memory budget05
  1. Trainable adapters

    LoRA and QLoRA train only these small matrices; full fine-tuning has no separate adapters.

  2. Optimizer state

    Kept for every trainable parameter, so it is small for adapters and large for full fine-tuning.

  3. Gradients and activations

    Grow with batch size and sequence length; gradient checkpointing trades compute for memory.

  4. Frozen base weights

    Held at training precision in LoRA, quantized to low precision in QLoRA, and fully trainable in full fine-tuning.

  5. Accelerator memory budget

    What the run must fit into, across one or several devices.

Conceptual view of the main consumers of accelerator memory in a fine-tuning run. Proportions vary with model size, sequence length and batch size.

LoRA, QLoRA and full fine-tuning compared

ColdAI's fine-tuning practice works with all three methods, on Llama, Mistral and custom architectures1. The choice between them should follow evidence like this.

CriterionLoRAQLoRAFull fine-tuning
Weights updatedAdapters only; base frozenAdapters only; base frozen and quantizedEvery weight
Memory pressureModerate: full-precision base plus small optimizer stateLowest: quantized base plus small optimizer stateHighest: optimizer state for every parameter
Training speedFastSomewhat slower, because weights are dequantized on the flySlowest per step, and needs the most hardware
Learning capacityEnough for most format, style and task behaviorClose to LoRA in the original paper's tests; verify on your taskHighest, for large shifts in domain or language
Forgetting riskLowerLowerHigher; needs regression testing and often mixed-in general data
Artifact to storeA small adapter tied to one base versionA small adapter tied to one quantized baseA full copy of the model
Best fitThe default first attemptLarge base models on limited hardwareEvidence that adapters plateau below the target

QLoRA's authors reported matching 16-bit full fine-tuning performance on their benchmarks while fine-tuning a 65B-parameter model on a single 48GB GPU3; your task needs its own comparison.

Hyperparameters to test first

Target modules often matter more than rank. The QLoRA authors found that applying adapters to all linear layers of each transformer block, not only the attention projections, was needed to match full fine-tuning3. Start there unless memory forces a narrower choice.

Rank and alpha set capacity and scale. Begin with a modest rank, keep alpha proportional to it so the effective update scale stays stable, and raise rank only if held-out results improve. With rank-stabilized scaling, higher ranks behave more predictably4.

Learning rate and epochs cause most failed runs. Adapters are usually trained with higher learning rates than full fine-tuning, but sweep a few values rather than copying one. Train for few epochs and stop when held-out loss or task scores stop improving; continuing usually memorizes the training set.

Formatting is the silent one: train with the exact chat template and special tokens the base model expects at inference, and mask the loss on prompt tokens if only the responses should be learned.

Regressions and forgetting after fine-tuning

General abilities degrade

Early signalThe tuned model handles the target task but fails simple requests the base model managed.

MitigationKeep a regression suite of general tasks. Comparisons of LoRA and full fine-tuning found LoRA forgets less of the base model's abilities but also learns less on demanding target domains5.

Safety behavior weakens

Early signalThe tuned model complies with requests the base model refused.

MitigationRe-run safety evaluations after every training run; research has shown that fine-tuning can erode safety alignment even on benign data6.

The model memorizes its training set

Early signalTraining scores keep rising while held-out scores flatten or fall, or outputs repeat training examples verbatim.

MitigationDeduplicate data, stop early, and probe for verbatim reproduction of sensitive examples.

Evaluation looks better than reality

Early signalHeld-out examples share documents, templates or customers with the training set.

MitigationSplit by source and period, as described in building a training dataset.

Serving: merge the adapter or hot-swap many

  • If

    You have one adapter and latency matters most.

    Then

    Merge it into the base weights and serve a standalone model4.

    A merged model adds no inference latency, at the cost of storing a full copy.

  • If

    Many tasks, teams or users each need their own behavior on one base model.

    Then

    Keep adapters separate and load them per request on a shared base.

    Serving systems have been shown to handle thousands of concurrent adapters on one base model7.

  • If

    The adapter was trained with QLoRA and you want to merge it.

    Then

    Load the base at higher precision, merge, and re-evaluate the merged model before release.

    Merging into quantized weights adds rounding error that can shift behavior.

  • If

    Fast rollback matters more than peak throughput.

    Then

    Serve versioned adapters, so rollback means switching back to the previous adapter.

    The base stays untouched, so the previous behavior is always one switch away.

When per-user adapters are worth evaluating

Per-user behavior is the clearest case for many adapters on one base. LinkTuned, one of ColdAI's case studies, is described as using a custom AI agent fine-tuned on each user's niche data, writing style and audience engagement patterns8. A design with that shape is where keeping small, separately versioned adapters over a shared base is worth testing against prompting with per-user examples.

License and data checks for an open-weight base model

0 of 6 checked

Questions and answers

Is DPO an alternative to LoRA?

No, they answer different questions. Direct Preference Optimization is a training objective that learns from pairs of preferred and rejected responses without a separate reward model10. LoRA is a way of choosing which parameters to train. You can run DPO on LoRA adapters, typically after a supervised fine-tuning stage has taught the basic task format.

Can a LoRA adapter be moved to a different base model?

Not reliably. An adapter encodes a change relative to specific base weights, so it belongs to the exact checkpoint it was trained on. A new version of the same family, or a differently quantized copy, needs the adapter retrained and re-evaluated. Keeping the training data and configuration versioned makes that retraining routine.

Does QLoRA reduce quality compared with LoRA?

In the QLoRA paper's experiments, quality matched full-precision fine-tuning on the benchmarks reported3, but results on your task may differ. Run both methods on a small slice of your data with the same evaluation set if you have the hardware; if you do not, QLoRA is usually the practical way to tune a larger base model.

How should a fine-tuned model be evaluated after training?

Compare it with the untuned base model and the best prompting baseline on the same frozen held-out set, scored the same way. Add a regression suite for general abilities, a safety evaluation, and latency and memory measurements on the serving setup. Release only if the gain holds on unseen inputs without unacceptable regressions.

Sources

  1. LLM Fine-tuning: methods and model families supported — ColdAI
  2. LoRA: Low-Rank Adaptation of Large Language Models (Hu et al.) — arXiv · checked 10 October 2026
  3. QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al.) — arXiv · checked 10 October 2026
  4. PEFT documentation: LoRA conceptual guide — Hugging Face · checked 10 October 2026
  5. LoRA Learns Less and Forgets Less (Biderman et al.) — arXiv · checked 10 October 2026
  6. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (Qi et al.) — arXiv · checked 10 October 2026
  7. S-LoRA: Serving Thousands of Concurrent LoRA Adapters (Sheng et al.) — arXiv · checked 10 October 2026
  8. LinkTuned: AI post automation platform — ColdAI
  9. Llama 3.1 Community License Agreement — Meta · checked 10 October 2026
  10. Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al.) — arXiv · checked 10 October 2026

More in LLM Fine-tuning

Back to LLM Fine-tuning

Next step

Choose a fine-tuning method against your hardware and task

Tell us the base model you are considering, the hardware available and the behavior you need. We will suggest which method to test first and the evaluation that should decide between them.

Plan a fine-tuning run