Deep diveEdge AI & IoT
Model quantization, pruning and distillation for constrained edge devices
Model quantization for edge devices stores weights and activations in low-precision integers so a model needs less memory and runs on integer hardware; pruning removes parameters the model can do without; distillation trains a smaller model to imitate a larger one. They work best combined, in a deliberate order, and compiled for a specific runtime. The proof is always measured on the target device, never on a workstation.
On this page
- Why the device, not the dataset, sets the design problem
- Compression terms an embedded ML engineer needs
- What each technique changes, and what it costs you
- Choosing between post-training and quantization-aware training
- An order of operations that avoids wasted compression work
- Runtimes and toolchains matched to edge hardware classes
- Evidence to collect on the target device before sign-off
- Questions and answers
- Sources
Why the device, not the dataset, sets the design problem
A model trained on a GPU server assumes plenty of memory, floating-point arithmetic and a stable power supply. An edge device offers none of those reliably. A microcontroller may have less working memory than the model's weights occupy, an application processor may share a small thermal envelope with everything else on the board, and a battery-powered sensor pays for every inference in energy.
So compression is not a final polishing step. The questions that matter come before training is finished: how much RAM is free at inference time once the firmware is loaded, whether the chip has an integer accelerator or neural processing unit (NPU), which operators that accelerator supports, and how long the device may take to decide before the decision is useless. ColdAI's edge practice optimizes models through quantization, pruning and knowledge distillation to run on ARM processors, FPGAs and custom silicon6, and every one of those targets rewards a different mix of the three techniques.
Compression terms an embedded ML engineer needs
- Scale and zero point
- The two parameters that map a range of real values onto integers. A tensor quantized to INT8 stores small integers plus the scale and zero point needed to recover approximate real values.
- Calibration data
- A small, representative set of inputs run through the model to observe activation ranges before fixing quantization parameters. Static quantization in ONNX Runtime, for example, depends on it and offers MinMax, Entropy and Percentile calibration methods2.
- Post-training quantization (PTQ)
- Converting an already-trained model to lower precision without retraining. Fast and cheap; accuracy loss depends on how sensitive the model is to rounding.
- Quantization-aware training (QAT)
- Fine-tuning with simulated quantization in the forward pass so the weights adapt to the rounding they will meet on the device. Slower, needs training data and pipeline, usually recovers accuracy that PTQ loses.
- Per-channel quantization
- Giving each output channel of a layer its own scale rather than one scale per tensor. It often preserves accuracy in convolutional layers whose channels have very different ranges.
- Structured pruning
- Removing whole filters, channels, attention heads or layers, so the remaining network is smaller and dense. Any hardware runs it faster because there is simply less arithmetic.
- Unstructured sparsity
- Zeroing individual weights wherever they matter least. The file compresses well, but the dense shape is unchanged, so speed only improves with sparse kernels or hardware built to skip zeros.
- Teacher and student
- In knowledge distillation, the large trained model (teacher) supplies soft output probabilities that a smaller model (student) learns to match, alongside the true labels5.
- Operator fallback
- When an accelerator cannot run an operator, the runtime hands it back to the CPU. A single unsupported layer can erase the gains of quantizing everything else.
What each technique changes, and what it costs you
| Question | Quantization | Pruning | Distillation |
|---|---|---|---|
| What shrinks | Bytes per weight and activation | Number of parameters or channels | The whole architecture |
| Retraining needed | No for PTQ; yes for QAT | Usually a fine-tuning pass after each pruning round | Yes: the student is trained from scratch or from a small base |
| Where speed comes from | Integer units, NPUs and lower memory traffic | Less arithmetic, if pruning is structured | A smaller network designed for the target |
| Typical failure | Outlier activations clip and accuracy drops on rare classes | Unstructured sparsity that saves storage but not time | A student that matches the teacher on common cases only |
| Effort | Hours to days | Days, with tuning of how much to remove | Weeks, because it is a training project |
Effort ranges are qualitative and depend heavily on model size and how mature the training pipeline is.
Choosing between post-training and quantization-aware training
Begin with PTQ. Static PTQ, where activation ranges are fixed from calibration data, suits convolutional networks well and is what most integer accelerators expect; ONNX Runtime's documentation recommends dynamic quantization for recurrent and transformer models and static quantization for CNNs2. The calibration set should look like field data, including night-time images, cold-start readings or noisy channels, because ranges learned on clean lab data clip the very inputs the device will struggle with.
Move to QAT when PTQ costs more accuracy than the decision can tolerate, or when you need to go below INT8, for example to INT4 weights. Runtimes differ here: TensorRT's Model Optimizer offers both PTQ and QAT3, while ONNX Runtime does not retrain and expects you to run QAT in the original framework and convert the result back2. Before you blame the method, check which layers lose the most: keeping a sensitive first or last layer in higher precision is often enough.
An order of operations that avoids wasted compression work
- Fix target and runtime
Chip, free memory, supported operators and latency budget written down first.
- Baseline on the device
Run the float model, or the nearest that fits, to measure where time and memory go.
- Distill if far too big
Only when the architecture is several sizes too large for the target.
- Prune structurally
Remove channels or blocks the profile shows are expensive, then fine-tune.
- Quantize: PTQ, then QAT
Calibrate on field-like data; fall back to QAT or mixed precision if needed.
- Compile and profile
Convert for the runtime and check for operators falling back to the CPU.
- Validate on device
Accuracy, latency, peak memory and energy on held-out field data.
Runtimes and toolchains matched to edge hardware classes
Name the runtime early: it determines which quantization formats, operators and accelerators are actually available.
| Hardware class | Commonly used runtime | What to check before compressing |
|---|---|---|
| Microcontrollers | LiteRT, the successor to TensorFlow Lite, which documents microcontroller deployment1 | Whether the model and its working buffers fit in on-chip RAM and flash with the firmware |
| Phones and NPU-equipped SoCs | LiteRT with NPU delegates1, or ONNX Runtime | Which operators and integer formats the vendor NPU supports, and what falls back to CPU |
| GPU edge modules | TensorRT, which lists the Jetson edge platform as a target3 | Whether FP16 alone meets the budget before attempting INT8 calibration |
| Intel CPUs, GPUs and NPUs | OpenVINO, whose NNCF provides PTQ, QAT and filter pruning4 | Which device plugin will run the model and how it handles mixed precision |
| FPGAs and custom silicon | Vendor-specific compilers | Fixed-point formats, layer support and how long a recompile takes when the model changes |
Toolchains change quickly. Treat this as a starting map and confirm support for your exact operators and versions.
Evidence to collect on the target device before sign-off
Questions and answers
How much accuracy should we expect to lose when quantizing to INT8?
It depends on the model and data, so measure rather than assume. Many convolutional models lose little with good calibration, while models with outlier activations, very small models and rare classes can lose more. Compare per-class results against the float model on field-like data, and treat an unexplained drop on a rare but important class as a release blocker even if overall accuracy looks fine.
Where does calibration data for post-training quantization come from?
Take a few hundred to a few thousand unlabeled samples that represent real operating conditions, drawn from the same pipeline that feeds the device. Labels are not needed for calibration, but coverage is: include edge conditions such as low light, temperature extremes or sensor noise, otherwise the learned ranges will clip exactly the inputs that matter.
Should we choose the hardware or the model first?
Shortlist the hardware first, because the runtime, supported operators and memory decide what compression is possible. Then choose or design a model family that the shortlisted chips run well. Committing to a model architecture before knowing which operators the accelerator supports is the most common cause of late, expensive rework.
Can pruning and quantization be combined with distillation?
Yes, and they often are. A typical sequence distills a large model into a compact student, prunes structurally where profiling shows cost, and then quantizes. Each step changes the model, so re-validate after each one rather than only at the end, and keep the float version of every stage so you can find which step caused a regression.
Sources
- LiteRT overview — Google AI Edge · checked 10 October 2026
- Quantize ONNX models — ONNX Runtime · checked 10 October 2026
- NVIDIA TensorRT — NVIDIA · checked 10 October 2026
- Model optimization guide — OpenVINO documentation · checked 10 October 2026
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals, Dean) — arXiv · checked 10 October 2026
- Edge AI & IoT — ColdAI