SCALECUA : Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
Abstract
Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, applying this paradigm is severely bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs by combining verifiable task synthesis with efficient online RL. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks for online RL. To maximize sample efficiency over this scaled pool, we propose Frontier Sampling, which dynamically tracks the model's per-task capability and allocates rollouts to tasks at the current learning frontier. On the training side, we further design Visual Context Segmentation, which keeps the recent visual context within a bounded window to balance rollout and training engine pressure on long-horizon trajectories, yielding a 2.83× training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Source code: