Trust, but Verify: Reliable Long-Horizon Mobile Agents Fully Solve AndroidWorld
Abstract
As computer-use agents move into real-world deployment, their dominant failure mode is quiet: the agent's belief about what it did diverges from what actually happened, and the error compounds over a long horizon. We present Minitap,\footnote{The system is available as open-source software; repository link withheld for anonymity.} a multi-agent system for mobile devices that treats reliability as a design constraint rather than an emergent property---and, as a result, achieves 100\% success on the AndroidWorld benchmark, the first system to fully solve all 116 tasks and to surpass human performance (80\%). We first analyze why a monolithic single-agent baseline fails (64\%), identifying six failure modes across context management, execution reliability, and error recovery. Minitap answers each with a trust mechanism: \emph{discretion is bounded}, by decomposing the agent into six roles with narrow mandates and focused contexts; \emph{action is verified}, by a deterministic layer that re-reads device state after fragile operations, so the agent cannot silently believe it acted when it did not; and \emph{behavior is overseen}, by a watchdog that analyzes decision history to detect loops and force strategy changes. Ablations attribute +21 points to the decomposition, +15 to verified execution, and +9 to the watchdog. A nine-configuration study shows reliability concentrates where judgment does: downgrading the single decision seat collapses success to 11\%, while budget models everywhere else match all-frontier reliability at 32\% lower cost. We close with implications for building computer-use agents that can be trusted in the wild.