Control OSWorld: An AI Control Environment for GUI Computer Use Agents
Abstract
AI agents that operate a computer through its graphical user interface (GUI) are being deployed widely. AI control studies how to prevent an AI system from causing harm even if it is misaligned and actively trying to do so. Most control research to date has focused on coding agents, leaving computer use largely unexplored. We introduce Control OSWorld, a control evaluation that pairs 318 tasks from OSWorld with 81 harmful side tasks (e.g., exfiltrating a private file) that an agent must complete without being caught. We use Control OSWorld to create and study control monitors that detect malicious agents interacting with a GUI. We find that a weaker monitor can reliably distinguish honest from malicious trajectories produced by a more capable agent when it sees the full trajectory (97% recall at 3% false positive rate). However, when the monitor must judge each step before it executes, recall drops at low false positive rates since there is more limited context to judge actions against. Monitor performance also depends on what the monitor can see. Access to the agent's visible text is critical, while screenshots provide little uplift.