PéPITE: Position-Preserving Image Token Exclusion for Efficient GUI Grounding
Abstract
GUI-grounding models, which map a natural-language instruction and a screenshot to a click position, are a core component of computer-use agents, and image tokens account for most of their serving cost: prefill compute, KV-cache memory, and image-embedding-cache capacity. We introduce Position-Preserving Image Token Exclusion (PéPITE), a training-free method that scores each image token's redundancy by mean pairwise cosine similarity of the vision-encoder embeddings and drops the most redundant fraction ρ. Existing token-reduction methods renumber or merge the surviving positions, which discards the spatial identity that localization depends on, and often condition selection on the instruction, which prevents caching; PéPITE keeps each survivor's original 2D position and scores the image alone, so a pruned image is encoded once and reused across requests. We evaluate seven open-source grounding models on ScreenSpot-v2 and ScreenSpot-Pro, against matched-budget random-pruning and downscaling baselines: at ρ=0.5, the median accuracy drop is 2.2 points on ScreenSpot-v2 and 4.0 on ScreenSpot-Pro, and serving throughput rises by up to 26% where language-model prefill dominates. The savings in KV-cache and embedding-cache memory matter most where memory is scarce: the consumer and edge hardware where 3B-9B grounding models run locally. Mechanistic analyses of the hybrid-attention Qwen3.5 architecture support the approach: the model's own recurrent update gates already suppress high-redundancy tokens, and pruning does not reduce the attention mass on the grounding target. We release a vLLM implementation and the full experiment code.