Decision-Aware Rounding for Low-Bit LLMs
Abstract
Low-bit post-training quantization (PTQ) makes small language models cheaper to store and serve, but standard reconstruction objectives do not indicate whether a rounding choice will change the model's answer. We introduce Decision-Aware Rounding (DAR), which guides adaptive rounding using the predicted first-order change in pairwise sequence-answer margins. For each calibration example answered correctly by the full-precision (FP) model, DAR projects the realized quantization perturbation onto the FP gradient separating the correct answer from each competitor. A one-sided objective penalizes predicted margin reductions, while DAR-Sym penalizes predicted margin changes in either direction. In controlled 3-bit experiments on the final two transformer blocks, the projection closely tracks the margin changes produced by hard quantization. DAR-Sym reduces harmful flips by more than threefold relative to matched reconstruction-only rounding on Qwen2.5-0.5B-Instruct, and a one-sided variant also improves decision preservation on SmolLM2-360M-Instruct. The results support answer-margin distortion as a useful task-aware signal for trustworthy low-bit rounding, while the present evidence remains local rather than establishing full-model deployment performance.