When Do Agent Safety Benchmarks Test Runtime Enforcement? Exposure-Aware Evaluation of Tool-Using Agents
Abstract
Tool-using agents increasingly place effectful actions behind runtime guards. End-to-end safety outcomes characterize the composed system under a tested distribution; empirical support for a conditional claim about a particular guard obligation additionally requires trajectories that reach and violate that obligation. We operationalize boundary-specific obligation exposure as this claim-support condition. In a frozen 1,044-episode hosted evaluation, 432 adversarial episodes produce zero unsafe executions, and N(A_B6)=0 for final state-bound authorization, leaving its finite-sample conditional-containment estimator undefined. A controlled surface localizes all 24 declared violations, admits all 6 legal controls, and detects all 6/6 injected B6 regressions; trusted-context perturbation also exposed a development-time source-of-truth defect. A separate 576-call hosted evaluation yields two B6-target exposures, both denied. A planner-oriented 400-episode unsafe-realization campaign yields zero unsafe proposals; direct token/state perturbation is reserved for the controlled surface. The exposure-aware protocol preserves first-boundary attribution and separates current-distribution system evidence from controlled guard-regression evidence. Zero observed unsafe executions can therefore coexist with an empty conditional evidence set for a relied-upon downstream safeguard.