Quantization-Aware Distillation for 4-Bit On-Device Language Models
Abstract
Quantization enables language models to run on memory- and compute-constrained devices, but low-bit formats can incur substantial quality loss. We study whether quantization-aware distillation (QAD) can recover this loss by training explicitly against the quantizer used at inference time. We study the 230M, 350M, and 1.2B-Instruct checkpoints from LFM2.5, an open-weight hybrid family designed for low-memory, low-latency on-device inference. Our target is Q40, a blockwise four-bit integer format implemented in llama.cpp, and we match its block structure, scale dtype, per-tensor type assignment, and activation path used by its CPU kernels. Across seven benchmarks, QAD raises Q40 from 86.8–92.6% of the BF16 aggregate score to 96.5–97.7%, recovering 64.9–73.4% of the quality lost to post-training quantization (PTQ). At 350M, matched supervised and quantization-aware fine-tuning controls on the same data fall below PTQ, so the gain is not explained by additional training. The QAD artifact shares PTQ Q40's file format, size, and kernels, so the gain carries no deployment cost. We find that recovery does not come from better alignment to the quantized weight grid: QAD increases weight-reconstruction error on the Q40 grid by 4.3–7.4% while reducing KL divergence to the BF16 teacher by 40.3–53.6%, adapting the quantized computation rather than moving weights closer to the quantization grid. The adaptations are also partly specific to the trained quantization map. When exported through a second, structurally similar four-bit format, checkpoints lose 0.9–1.6 points of their gains. Our results suggest treating the deployment quantizer as part of the training target for low-bit on-device language models.