When Does an Agent Benchmark Test the Verifier? Exposure-Aware Evaluation of Runtime Enforcement
Abstract
An agent benchmark can record zero unsafe executions even when the runtime-verification obligation under study receives no exposure. We study this attribution problem for tool-using agents with a stateful runtime verifier. Our evaluation separates boundary-specific verification exposure from conditional containment—denial given exposure—and adds exposure-controlled falsification: each predeclared violation must reach its earliest violated boundary and trigger the declared rejection, with matched legal controls preserving admissible behavior. In a 1,044-episode hosted evaluation, 432 adversarial episodes yield zero unsafe executions and zero proposals that both reach the final state-bound authorization boundary (B6) and violate its contract. Conditional containment at B6 is therefore undefined on the primary denominator. Controlled evaluation localizes all 24 violations, admits all 6 legal controls, and matches every predeclared outcome in the policy/state and token suites. A mutation-sensitivity study removes one of six B6 invariants spanning intent, policy, authority, state, expiry, and replay; the unchanged token suite detects all 6/6 injected regressions. A separately hosted set realizes two proposals that reach and violate B6, both denied. For evolving agents, we propose a two-axis design that pairs distribution-specific system evaluation with a fixed controlled verifier-regression suite.