Gate the Enforceable, Record the Rest: A Pre-Execution Guard Measured at Both Ends
Abstract
In a six-arm ablation holding rules, tasks and model fixed, a self-service override runs at least 48.0% in the arm whose refusal names it, and exactly 0.0% across the 416 decisions of the four arms that do not. No multiplicity correction reaches a separation that exact. A guardrail is not a classifier but an intervention, and precision and recall see neither of its ends: what the agent may attempt, and what it does next. Explaining the refusal is the weaker claim. Rewriting rises from 40.8% to 66.7% when the refusal explains itself, but a length-matched placebo with no reason in it reaches 51.4%, taking 10.6 of those 25.9 points, and an arm stating the reason in full while instructing the agent to stop rewrites at 41.0%, below the placebo. A reason alone does not produce the lift. What separates the two arms is the absence of an instruction not to act on it. Three models answer a byte-identical refusal at 48.0%, 32.0% and 7.3%. The action space fixes what a guard can decide before any rule is written. Within 7,034 tool calls in 644 sessions where one agent chose freely between a shell and structured file tools, one 21-rule guard refuses 19.3% of mutating shell commands against 6.3% of mutating structured calls (+13.1 points, session-clustered 95% CI [+8.6, +17.6]) on 22 structured events from one rule. Expressibility is the mechanism, and a second monitor holding harm class fixed reverses the ordering. Measuring the guard perturbs it: the fault rate tracked disk pressure on the host, 6.5% -> 9.6% and back to 2.9% on a prediction registered before the fix. Developers write action-local rules 4.1-6.1% of the time across 166,193 statements. One class escapes all three. Undecidable before execution and irreversible after, it is 5.2% [3.8, 7.1] of one public corpus, and 4.2% [2.9, 6.1] of our own sessions took an action reaching outside their sandbox.