GateArena: Environment-Enforced Confirmation for Irreversible Agent Actions
Abstract
LLM agents are increasingly given tools whose effects cannot be undone. The prevailing safeguard is an instruction in the system prompt directing the agent to confirm with the user first, which makes the safety property contingent on model behaviour. We measure that contingency: given a byte-identical instruction, tool set, task distribution and environment seeds, three current backbones attempted to call a withheld confirmation tool in 38%, 48% and under 1% of episodes, with no adversarial input involved. Binding an authorization to the specific call it permits is the classical remedy, and we apply it here rather than propose it. A confirmation returns a warrant, a token that binds the approved action to its exact parameter values and is consumed by the execution it authorises, and the execution tool rejects any irreversible call whose warrant is missing, already spent, or bound to other values. To evaluate it we build GateArena, a dual-control environment for irreversible tool calls with two irreversible action types, a controllable rate at which stated requests deviate from the user's intent, and harm adjudicated against that intent rather than against the request. Across three backbones and 1,880 episodes, ungated execution reaches 37.5-60% irreversible harm at a 60% deviation rate while the gate holds harm at 2.5%. Holding the instruction and tool set fixed and moving only the check into the handler takes protocol-violating executions from 91/400 to 0/400, and a scenario ablation separates what the binding adds over a boolean confirmation flag. Code, environment and episode logs are released.