Sound Plans, Forged Values: Cross-Channel Constraint Integrity for Computer-Use Agents
Abstract
Computer Use Agents (CUAs) directly operate graphical interfaces and invoke Model Context Protocol (MCP) tools exposed by website vendors. This makes indirect prompt injection an immediate attack vector: every rendered page and every tool are attack injection points to the agent's control flow. The dual-LLM architecture is proposed as a system-level defense against this. A privileged planner (P-LLM) constructs a plan from the user's query before any untrusted data is seen, while a quarantined model (Q-LLM) processes data but is structurally prevented from influencing the control flow. Since the plan is fixed ahead-of-time, injected instructions cannot redirect the agent. However, this guarantee falls short for CUAs, where the control flow is inherently data-dependent: a CUA cannot know a page's contents, until it opens the page. Plans must therefore contain branches, and branches themselves are another injection surface, and this is known as branch steering attacks: the adversary injects no new instruction, but supplies data that drives an already-approved plan down an attacker-chosen path. To this end, we study this branch steering attacks systematically and propose a new dual-LLM structure to defend against it. We first present STEER-Bench, a suite of 101 end-to-end CUA tasks spanning 9 domains, and show that branch steering succeeds at 80% against standard CUAs and 77% against vanilla dual-LLM CUAs, giving the first systematic characterization of this attack class for CUAs. Then, we introduce COBRA, a novel dual-LLM CUA system that emits not only a trusted plan, but also an ahead-of-time enforcement contract over HTTP requests and MCP tool calls, binding each branch to the concrete effects the planner allows. On STEER-Bench, COBRA reduces attack success rate to 0% while retaining 88% of clean-task utility.