Post-Training Quantization with Gradient-Projected Fisher Approximation for Vision Transformers
Abstract
Vision Transformers (ViTs) demonstrate strong performance, but their substantial memory footprint and computational overhead necessitate efficient compression techniques such as post-training quantization (PTQ). However, pushing ViTs to low-bit precision incurs significant accuracy degradation. Recent PTQ methods adopt block reconstruction-based optimization with curvature-aware objectives, typically implemented using Fisher-based approximations. These approaches rely on explicit curvature modeling based on diagonal or structured approximations, which still fail to capture cross-dimensional interactions. Moreover, they often involve matrix inversion, leading to numerical instability under ill-conditioned settings. To address these limitations, we propose Gradient-Projected Fisher Approximation for Quantization (GPFA-Q), a block reconstruction-based PTQ framework that avoids explicit curvature matrix construction while capturing off-diagonal interactions. First, we introduce Gradient-Projected Reconstruction (GPR), a reconstruction objective that captures cross-dimensional interactions without explicitly constructing curvature matrices. To further support GPR, we integrate Soft Grid Rounding (SGR), which reduces the mismatch between continuous reconstruction and discrete inference. Extensive experiments demonstrate that our GPFA-Q achieves the state-of-the-art performance in low-bit quantization across diverse vision tasks.