GazeZoom: Replacing Crop Tools with Attention in Small VLM Agents
Abstract
Multimodal agents often rely on high-resolution inputs or selective inspection through crop tools; generating crop requests adds output tokens to an already lengthy transcription. We introduce GazeZoom, an attention-guided localization method for recovering fine details across complex images. We first apply question-conditioned localization to static document QA, then extend it to Live GazeZoom: attention from a provisional token span selects a higher-resolution detail for regenerating that span, while preserving the overview and accepted text. On manga OCR with Qwen3-VL-8B, Live GazeZoom raises content F1 from 48.1\% to 61.1\% over a crop-tool agent at the same overview and detail image-token budgets, while producing fewer generated tokens on average. These results connect the model's changing spatial focus to improved text recovery without explicit crop requests. Weaker results on OmniDocBench motivate coverage-aware crop policies; cross-call cache reuse remains a direction for reducing inference cost.