Rubrics Before Rollouts: Grounded Verifiers for Test-Time Scaling of Computer-Use Agents
Abstract
Computer-use Agents (CUA) are increasingly advancing toward tackling complex, open-ended desktop tasks. However, reliable verification is particularly difficult to scale in these GUI settings: for example, the benchmark OSWorld-Verified uses function-based evaluators that require hand-crafted Python scripts per task. This bottlenecks exploring inference-time gains through Test-Time Scaling (TTS) for such agents. To address this, we explore Environment-Grounded Rubrics for CUA, a framework in which a rubric agent receives a task description and access to the initialized virtual machine, and through exploration prior to any model rollouts, synthesizes a context-grounded rubric checklist across five axes: Task Requirements (TR), System State Outcomes (SS), Visible UI Evidence (VE), Safety and Non-interference (SN), and Failure Detection (FD). Candidate trajectories are then scored against the rubric without requiring access to ground-truth evaluation scripts. On OSWorld-Verified under parallel TTS evaluation with K=16 rollouts, Environment-Grounded Rubrics for CUA achieves a Best@16 score of 70.2% with Claude Sonnet 4.6, with at least a 2.6 percentage-point gain (and 15% improvement in the coverage of the Oracle-Random gap) over the strongest non-agentic, non-rubric verification baseline. Crucially, Environment-Grounded Rubrics for CUA achieves this while being 5.5× cheaper and 3.4× faster than the strongest fully agentic alternative, representing the best cost-capability tradeoff among all verifiers evaluated. Our ablations show that this rubric paradigm is robust to the choice of both rubric generation model and judge model, with all frontier models evaluated mostly falling within 0.5 pp of each other. These findings position Environment-Grounded Rubrics for CUA as a practical path toward efficient verification for the test-time scaling of the next generation of computer-use agents.