The Denial Illusion: Measuring Safety Failures After Agent Replanning
Abstract
A rejected tool call does not establish that its forbidden effect was prevented: an agent can continue through another tool, split the operation across calls, or reach the same state through an intermediate resource. We introduce denial-induced route-around and RouteAroundBench, a controlled post-denial benchmark with executable tools and deterministic terminal-state oracles. The study compares three purpose-built controls (exact-call locking, tuple-level locking, and a full-trace language-model guard) across three locally served open-weight models. We also present DenyLock, a training-free reference monitor that retains the structured rule responsible for a denial and checks every tentative state transition before atomic commit. Across 120 unsafe tasks and 60 paired benign tasks, exact-call locking permits the forbidden effect in 56.9% of post-denial trials; a trace-aware guard lowers this rate to 20.8%. Under the benchmark's complete mediation assumptions, DenyLock prevents every forbidden transition while retaining 91.1% benign recovery. These results distinguish a locally correct denial from trajectory-level policy preservation. Code, the RouteAroundBench specification, and a deterministic end-to-end validation script are available at https://anonymous.4open.science/r/denylock-artifact-3CBA.