Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents
Abstract
Causal action verifiers gate an agent's state-changing tool calls by checking whether the proposed intervention is identifiable against a committed action-state graph, and issue a certificate carrying the identification argument and a one-sided lower confidence bound. A recently published verifier of this kind reports zero false executions across a confounded tool-use benchmark. We show that this guarantee is a property of the committed graph, not of the verifier, and we measure how fast it fails. Omitting a single bidirected edge, so the graph asserts no unmeasured confounding where some exists, takes the verifier from zero false executions to 15.3% at the benchmark's published confounding strength, with 91% of everything it chooses to execute being wrong and utility falling from +2.27 to +0.35; above that strength utility is negative and every execution is false. Reversing one arrowhead on a variable that genuinely exists, so a mediator is committed as a confounder, inverts the verifier completely: 48.9% false executions, zero correct executions, utility -1.36. Every one of these actions ships a valid certificate. We then show that safety is cheaply recoverable and utility is not. An attestation step that spends a bounded experiment on each observationally identified execution restores zero false executions at every confounding strength and is a no-op on honest graphs, yet it recovers none of the lost value: 97.1% of beneficial actions remain unexecuted with or without it. A gate that inspects proposed actions can only prevent wrongful action. Wrongful inaction produced by the same misspecification emits no certificate, produces no downstream signal, and is structurally invisible to it.