GPR: Grounding-Preserving Recovery for Quantized Vision-Language Models in Edge Deployments
Abstract
Vision-language models are increasingly being deployed in the real world on memory-constrained edge devices through quantization. However, performance on vision tasks that require precise spatial understanding can be degraded by standard quantization methods from 16-bit to 4-bit. We first find that quantized models can retain strong general ability on visual-question-answering and hallucination benchmarks (VQAv2 and POPE) while giving up substantial accuracy on referring expression comprehension (RefCOCO). On two tested models, Qwen3-VL-2B and LFM2.5-VL-1.6B, quantization to 4 bits causes a drop of 10.7 and 5.5 points on RefCOCO respectively. To address this, we propose Grounding-Preserving Recovery (GPR), a training-free method for restoring visual grounding in a model after quantization, by identifying where to allocate a small amount of additional memory. Our method locates the most valuable weights for grounding by selectively testing the recovery of individual decoder projections, restoring each from 4 bits to 8 bits and measuring how closely the model's coordinate outputs realign with the full-precision original. On both models, GPR recovers 69.2\% and 30.9\% of the lost grounding accuracy at only around 2\% of added memory, exceeding matched-memory allocations guided by general answer quality or modality balance, while closely maintaining general performance.