Concealment Collapse: Verification Externalities across Agent Handoffs
Tomoya Hoshino
Abstract
Adding verification checkpoints is a natural design choice for improving reliability, but this monotonicity breaks down when an agent being verified can strategically conceal evidence under incentives for reward hacking. We formulate verification-checkpoint placement on a workflow DAG as a strategic evidence-disclosure game with a handoff externality (evidence attenuation at handoffs). Introducing a second verifier can make the disclosure payoff non-concave, which can cause the optimal disclosure to switch discontinuously from full disclosure to full concealment. We call this phenomenon $emph{concealment collapse}$ and characterize its social loss through a closed-form Price of Anarchy. It increases with delegation depth in chains but is bounded by branching in trees. A controlled ground-truth oracle experiment consistently reproduces the theoretical predictions; on LiveCodeBench, a 14B verifier of the same class as the generator has detection capability near the floor, and the placement and disclosure effects cannot be identified. Placement and strategic disclosure shape reliability only when the verifier is sufficiently capable, and detection capability is a constraint that precedes placement: $emph{whether the verifier can detect}$ determines $emph{where to verify}$.
Chat is not available.
Successful Page Load