Towards Lightweight Visual Routing for Mobile Agents via Fine-Grained Screenshot Synthesis
Cheng-Che Lee ⋅ Chia-Pin Cheng ⋅ Hsu-Sheng Chang ⋅ CHIALIN CHOU
Abstract
Mobile agents rely heavily on visual routing—a perception-level decision task that determines the UI context of the current screen in order to decide which downstream tool to invoke. For such high-frequency, narrowly-scoped routing subtasks, continual reliance on general-purpose small vision-language models (sVLMs) may impose disproportionate latency and resource overhead on edge devices; we further find that mainstream sVLMs exhibit markedly limited zero-shot performance on this structured task. We therefore argue for replacing sVLMs with a task-specific, ultra-lightweight model (<5M parameters) as the router. Training such a model, however, requires labeled data, and real screenshots involving privacy-sensitive categories (SMS, instant messaging) are difficult to obtain at scale—this data scarcity, not model architecture, is the central challenge this paper addresses. We propose a diagnosis-driven, fine-grained screenshot synthesis methodology: unlike content-agnostic conventional augmentation, we first characterize the visual distribution gap between synthetic and real screenshots using multiple distribution-distance metrics, identifying conversation background style and modern system UI rendering as two dominant sources. We hypothesize that reducing these two gaps improves generalization, and accordingly design two targeted mechanisms—Bubble-blending and modern SVG UI rendering—both of which are shown to reduce the identified gaps. A three-stage progressive evaluation (learnability on synthetic data, transferability to unseen public apps, and generalizability to real screenshots) yields results consistent with this hypothesis, with the combined configuration giving the largest and most consistent gains across backbones, particularly on real data. As one measure of the resulting model’s practical utility, a 2.3M-parameter lightweight model trained entirely on synthetic and public data—without a single privacy-sensitive real screenshot—achieves 71.62% accuracy on a real-world test set, outperforming four general-purpose sVLMs with over 100$\times$ as many parameters evaluated under zero-shot prompting (e.g., PaliGemma 2 (3B), at 57.98%). These results support diagnosis-driven synthesis as a practical route to training task-specific, edge-deployable models under real-world data-access constraints.
Chat is not available.
Successful Page Load