Rethinking Gradient Approximation in Quantization: A Zeroth-Order Expectation Perspective
Abstract
As large language models (LLMs) continue to scale, quantization has become a key technique for efficient deployment. However, multi-bit quantization employs non-differentiable rounding, hindering gradient optimization. Existing methods rely on heuristic surrogate gradients (e.g., STE), which work empirically but lack a unified theory. To address this challenge, we propose the Zeroth-Order Expectation Gradient (ZOE-Grad), a zeroth-order expectation view of quantization surrogate gradients. Specifically, we establish an equivalence between the expectations of zeroth-order gradient estimators and surrogate gradients, providing a unified zeroth-order interpretation of widely used surrogates. Under this framework, STE corresponds to a degenerate, discontinuous perturbation that ignores the local geometry of quantization boundaries, limiting its expressiveness. In contrast, continuous perturbation-based surrogate gradients capture boundary-local information and produce smoother, more structured gradients. We further establish bounded-error convergence guarantees for the quantized objective. Through simulations under multiple perturbation distributions, we verify that the proposed framework accurately captures the relationship between the expectations of zeroth-order gradient estimators and surrogate gradients. Furthermore, extensive experiments on OPT-1.3B, OPT-6.7B, LLaMA-2-7B, and Qwen3-8B show that continuous ZOE-Grad surrogates consistently outperform STE across diverse LLM architectures.