PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
Jingxuan Wei ⋅ Xi Bai ⋅ Shan Liu ⋅ caijun jia ⋅ Zheng Sun ⋅ Xinglong Xu ⋅ Siyuan Li ⋅ Linzhuang Sun ⋅ Bihui Yu ⋅ Conghui He ⋅ Cheng Tan
Abstract
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving \emph{region-tolerant} paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions. Because geometric primitives carry ontological dependencies, a local coordinate error can induce cascading topological failures that distort downstream objects and invalidate the final construction. We identify this regime as \emph{precision-sensitive GUI tasks}, requiring point-level accuracy, geometry-aware verification, and robustness to dependency-driven error propagation. To benchmark it, we introduce \textsc{PAGE} Bench, with 4{,}906 problems and over 224K process-supervised, pixel-level GUI actions. We further propose \textsc{PAGER}, a topology-aware agent that decomposes construction into dependency-structured planning and pixel-level execution. Pixel-grounded supervised tuning establishes executable action grammar, while precision-aligned reinforcement learning mitigates rollout-induced exposure bias through state-conditioned geometric feedback. Experiments reveal a pronounced ``Semantic-Execution Gap'': general multimodal models can exceed 88\% Action Type Accuracy yet remain below 6\% Task Success. \textsc{PAGER} closes this gap, delivering $4.1\times$ higher Task Success than the strongest evaluated general baseline and raising Step Success Rate from below 9\% for GUI-specialized agents to over 62\%, establishing a new state of the art for point-precise GUI control.
Chat is not available.
Successful Page Load