Qwen-CUA: Native Computer Use for (almost) Everything
Abstract
Native computer use offers a general route to agents that can operate almost any software through the same interface available to people. Realizing this promise, however, requires more than visual grounding: an agent must sustain long-horizon state, acquire large amounts of costly interactive experience, and learn from outcomes that are sparse but reliably verifiable. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its agent scaffold expands the active visual history to 20 screenshots and folds older screenshots in fixed-size blocks, retaining recent visual evidence while preserving reusable prompt prefixes. To scale training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows in everyday and professional software. We optimize complete trajectories with verifiable rewards and long-horizon trajectory slicing, and develop the model through iterative training runs that use each resulting policy to refresh the supervised data mixture and calibrate the reinforcement-learning task distribution. Across eight computer-use benchmarks, Qwen-CUA consistently outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5 / 48.4 binary / partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion total parameters yields Qwen-CUA-Max, which further improves these scores to 87.6 and 21.2 / 53.3. Qwen-CUA also reduces attack success on RedTeamCUA from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, an internal browser deployment, and experiments combining native interaction with Bash further characterize its practical behavior. These results show that native computer use can serve as a broadly capable foundation, while scalable verifiable interaction and hybrid tool use are key to making it effective in practice.