ScaleBITS: Scalable Bitwidth Search for Hardware-Aligned Mixed-Precision LLMs
Xinlin Li ⋅ Timothy Chou ⋅ Joshua W Fromm ⋅ Zichang Liu ⋅ Yunjie Pan ⋅ Christina Fragouli
Abstract
Post-training weight quantization is crucial for reducing the deployment cost of large language models (LLMs), yet maintaining model quality in the ultra-low-bit regime ($\le$ 3 bits) remains challenging due to highly non-uniform weight sensitivity and the lack of principled precision allocation strategies. In this work, we show that the weight sensitivity is inherently dynamic and depends on the current quantized model, rather than the full-precision reference used in existing methods. Leveraging this observation, we introduce a sensitivity estimate that better captures the changes in marginal loss during progressive quantization. We further uncover a structured pattern in sensitivity distribution, revealing a bi-directional concentration across input and output channels, which enables hardware-aligned yet expressive partitioning via channel reordering. Building on these insights, we propose ScaleBITS, a scalable framework that combines structured partitioning and an efficient approximation to greedy allocation. Our approach enables fine-grained, global bitwidth search while preserving hardware efficiency. Experiments show that ScaleBITS significantly improves over uniform-precision quantization (up to +36%) and outperforms state-of-the-art sensitivity-aware baselines (up to +13%) in the ultra-low-bit regime, without adding runtime overhead.
Chat is not available.
Successful Page Load