Deep diveEdge AI & IoT

Model quantization, pruning and distillation for constrained edge devices

Model quantization for edge devices stores weights and activations in low-precision integers so a model needs less memory and runs on integer hardware; pruning removes parameters the model can do without; distillation trains a smaller model to imitate a larger one. They work best combined, in a deliberate order, and compiled for a specific runtime. The proof is always measured on the target device, never on a workstation.

Reviewed 8 min read

On this page
  1. Why the device, not the dataset, sets the design problem
  2. Compression terms an embedded ML engineer needs
  3. What each technique changes, and what it costs you
  4. Choosing between post-training and quantization-aware training
  5. An order of operations that avoids wasted compression work
  6. Runtimes and toolchains matched to edge hardware classes
  7. Evidence to collect on the target device before sign-off
  8. Questions and answers
  9. Sources

Why the device, not the dataset, sets the design problem

A model trained on a GPU server assumes plenty of memory, floating-point arithmetic and a stable power supply. An edge device offers none of those reliably. A microcontroller may have less working memory than the model's weights occupy, an application processor may share a small thermal envelope with everything else on the board, and a battery-powered sensor pays for every inference in energy.

So compression is not a final polishing step. The questions that matter come before training is finished: how much RAM is free at inference time once the firmware is loaded, whether the chip has an integer accelerator or neural processing unit (NPU), which operators that accelerator supports, and how long the device may take to decide before the decision is useless. ColdAI's edge practice optimizes models through quantization, pruning and knowledge distillation to run on ARM processors, FPGAs and custom silicon6, and every one of those targets rewards a different mix of the three techniques.

Compression terms an embedded ML engineer needs

Scale and zero point
The two parameters that map a range of real values onto integers. A tensor quantized to INT8 stores small integers plus the scale and zero point needed to recover approximate real values.
Calibration data
A small, representative set of inputs run through the model to observe activation ranges before fixing quantization parameters. Static quantization in ONNX Runtime, for example, depends on it and offers MinMax, Entropy and Percentile calibration methods2.
Post-training quantization (PTQ)
Converting an already-trained model to lower precision without retraining. Fast and cheap; accuracy loss depends on how sensitive the model is to rounding.
Quantization-aware training (QAT)
Fine-tuning with simulated quantization in the forward pass so the weights adapt to the rounding they will meet on the device. Slower, needs training data and pipeline, usually recovers accuracy that PTQ loses.
Per-channel quantization
Giving each output channel of a layer its own scale rather than one scale per tensor. It often preserves accuracy in convolutional layers whose channels have very different ranges.
Structured pruning
Removing whole filters, channels, attention heads or layers, so the remaining network is smaller and dense. Any hardware runs it faster because there is simply less arithmetic.
Unstructured sparsity
Zeroing individual weights wherever they matter least. The file compresses well, but the dense shape is unchanged, so speed only improves with sparse kernels or hardware built to skip zeros.
Teacher and student
In knowledge distillation, the large trained model (teacher) supplies soft output probabilities that a smaller model (student) learns to match, alongside the true labels5.
Operator fallback
When an accelerator cannot run an operator, the runtime hands it back to the CPU. A single unsupported layer can erase the gains of quantizing everything else.

What each technique changes, and what it costs you

QuestionQuantizationPruningDistillation
What shrinksBytes per weight and activationNumber of parameters or channelsThe whole architecture
Retraining neededNo for PTQ; yes for QATUsually a fine-tuning pass after each pruning roundYes: the student is trained from scratch or from a small base
Where speed comes fromInteger units, NPUs and lower memory trafficLess arithmetic, if pruning is structuredA smaller network designed for the target
Typical failureOutlier activations clip and accuracy drops on rare classesUnstructured sparsity that saves storage but not timeA student that matches the teacher on common cases only
EffortHours to daysDays, with tuning of how much to removeWeeks, because it is a training project

Effort ranges are qualitative and depend heavily on model size and how mature the training pipeline is.

Choosing between post-training and quantization-aware training

Begin with PTQ. Static PTQ, where activation ranges are fixed from calibration data, suits convolutional networks well and is what most integer accelerators expect; ONNX Runtime's documentation recommends dynamic quantization for recurrent and transformer models and static quantization for CNNs2. The calibration set should look like field data, including night-time images, cold-start readings or noisy channels, because ranges learned on clean lab data clip the very inputs the device will struggle with.

Move to QAT when PTQ costs more accuracy than the decision can tolerate, or when you need to go below INT8, for example to INT4 weights. Runtimes differ here: TensorRT's Model Optimizer offers both PTQ and QAT3, while ONNX Runtime does not retrain and expects you to run QAT in the original framework and convert the result back2. Before you blame the method, check which layers lose the most: keeping a sensitive first or last layer in higher precision is often enough.

An order of operations that avoids wasted compression work

far over budgetnear budgetmisses budget01Fix target and runtime02Baseline on the device03Distill if far too big04Prune structurally05Quantize: PTQ, thenQAT06Compile and profile07Validate on device
  1. Fix target and runtime

    Chip, free memory, supported operators and latency budget written down first.

  2. Baseline on the device

    Run the float model, or the nearest that fits, to measure where time and memory go.

  3. Distill if far too big

    Only when the architecture is several sizes too large for the target.

  4. Prune structurally

    Remove channels or blocks the profile shows are expensive, then fine-tune.

  5. Quantize: PTQ, then QAT

    Calibrate on field-like data; fall back to QAT or mixed precision if needed.

  6. Compile and profile

    Convert for the runtime and check for operators falling back to the CPU.

  7. Validate on device

    Accuracy, latency, peak memory and energy on held-out field data.

Conceptual sequence for compressing a model for an edge target. Steps are skipped when the model already fits; it is not a timeline or a measured result.

Runtimes and toolchains matched to edge hardware classes

Name the runtime early: it determines which quantization formats, operators and accelerators are actually available.

Hardware classCommonly used runtimeWhat to check before compressing
MicrocontrollersLiteRT, the successor to TensorFlow Lite, which documents microcontroller deployment1Whether the model and its working buffers fit in on-chip RAM and flash with the firmware
Phones and NPU-equipped SoCsLiteRT with NPU delegates1, or ONNX RuntimeWhich operators and integer formats the vendor NPU supports, and what falls back to CPU
GPU edge modulesTensorRT, which lists the Jetson edge platform as a target3Whether FP16 alone meets the budget before attempting INT8 calibration
Intel CPUs, GPUs and NPUsOpenVINO, whose NNCF provides PTQ, QAT and filter pruning4Which device plugin will run the model and how it handles mixed precision
FPGAs and custom siliconVendor-specific compilersFixed-point formats, layer support and how long a recompile takes when the model changes

Toolchains change quickly. Treat this as a starting map and confirm support for your exact operators and versions.

Evidence to collect on the target device before sign-off

0 of 6 checked

Questions and answers

How much accuracy should we expect to lose when quantizing to INT8?

It depends on the model and data, so measure rather than assume. Many convolutional models lose little with good calibration, while models with outlier activations, very small models and rare classes can lose more. Compare per-class results against the float model on field-like data, and treat an unexplained drop on a rare but important class as a release blocker even if overall accuracy looks fine.

Where does calibration data for post-training quantization come from?

Take a few hundred to a few thousand unlabeled samples that represent real operating conditions, drawn from the same pipeline that feeds the device. Labels are not needed for calibration, but coverage is: include edge conditions such as low light, temperature extremes or sensor noise, otherwise the learned ranges will clip exactly the inputs that matter.

Should we choose the hardware or the model first?

Shortlist the hardware first, because the runtime, supported operators and memory decide what compression is possible. Then choose or design a model family that the shortlisted chips run well. Committing to a model architecture before knowing which operators the accelerator supports is the most common cause of late, expensive rework.

Can pruning and quantization be combined with distillation?

Yes, and they often are. A typical sequence distills a large model into a compact student, prunes structurally where profiling shows cost, and then quantizes. Each step changes the model, so re-validate after each one rather than only at the end, and keep the float version of every stage so you can find which step caused a regression.

Sources

  1. LiteRT overview — Google AI Edge · checked 10 October 2026
  2. Quantize ONNX models — ONNX Runtime · checked 10 October 2026
  3. NVIDIA TensorRT — NVIDIA · checked 10 October 2026
  4. Model optimization guide — OpenVINO documentation · checked 10 October 2026
  5. Distilling the Knowledge in a Neural Network (Hinton, Vinyals, Dean) — arXiv · checked 10 October 2026
  6. Edge AI & IoT — ColdAI

More in Edge AI & IoT

Back to Edge AI & IoT

Next step

Send us your target board and the model that will not fit

Share the hardware, runtime, memory budget and a sample of field data. We will come back with a compression plan and the on-device tests that would prove it.

Discuss a compression plan