Self-Calibrated GUI Reward Model via Inverse Dynamic Modeling
Abstract
Developing generalist GUI agents capable of robust long-horizon execution remains a key challenge. While Reinforcement Learning (RL) offers a promising path for improvement, it is severely bottlenecked by the lack of reliable reward signals in open-ended environments. Ground-truth verification typically relies on expensive human annotation, which is prohibitively scarce at the scale required for training, while existing VLM-based judges frequently hallucinate success on superficial visual cues. This scarcity of trustworthy supervision fundamentally limits the scalability of current methods. In this work, we propose I-Judge, a novel framework for Self-Calibrated Reward Modeling. We leverage Inverse Dynamic Modeling (IDM) as a self-supervised training objective to learn the causal dynamics of GUI interactions from massive unlabeled trajectories, effectively bypassing the bottleneck of scarce human labels. We then introduce a runtime calibration mechanism that weights reward signals by the IDM's action consistency, filtering out spurious successes. Extensive end-to-end RL experiments on OSWorld, ScienceBoard and AndroidWorld demonstrate that our method significantly accelerates convergence and improves final agent performance. Notably, I-Judge demonstrates strong generalization to unseen domains. All the code, data and models will be made publicly available to foster further research.