Unifying Sparsity and Discreteness: One-Shot Pruning for Quantized LLMs via Discrete Optimization
Abstract
Pruning and quantization are two dominant techniques that effectively address the computational and storage burdens of large language model (LLM) inference on edge devices. Recently, one-shot pruning has gained particular attention for its ability to identify weight supports via optimization without retraining. However, integrating such methods seamlessly with quantization for further model lightweighting is non-trivial. The discreteness of quantization shatters their required variable continuity, inevitably collapsing this integration into a decoupled pruning-quantization pipeline with inherently suboptimal outcomes. To tackle this problem, we propose Quantization-aware One-shot Pruning (QOP), which directly optimizes the quantized weights under sparsity and discreteness constraints, thereby explicitly capturing the impact of quantization on pruning within its objective. Specifically, QOP generalizes the alternating direction method of multipliers to sparse-constrained discrete optimization, enabling the identification of the high-quality support and the update of quantized weights. Theoretically, via the well-conditioned approximation obtained by a slight perturbation, we establish convergence to a stationary feasible point and provide the convergence rate. Extensive experiments across various LLMs, sparsity levels, and quantization settings demonstrate that QOP consistently outperforms existing baselines in terms of both average accuracy and perplexity under weight quantization.