Watching the Watchers: Fault-Injection Evaluation of Trace-Based Verifiers for Multi-Agent Workflows
Abstract
Organizations deploying AI agents need evidence that the systems watching them can distinguish policy violations from ordinary work. When one model evaluates other agents, that verifier must itself be tested. In a simulated procurement and finance workflow, eight violations are injected across event-local evidence, relations spanning multiple events, and value claims requiring an authoritative source. Each violation forms a same-seed counterfactual pair: a violating trace and a matched benign trace that removes only the policy-defining relation. Deterministic rule-based verifiers are compared with an Agent4Agent verifier implemented using a large language model (LLM); an alarm receives credit only if it identifies a ground-truth violation event. Event-local rules achieve an area under the receiver operating characteristic curve (AUROC) of 0.95–0.98 but remain at chance on cross-event relations. Trace-level predicates recover three such relations at 0.975–0.978, while value validity remains unresolved without external grounding. Across 5,760 LLM evaluations covering three open-weight model families and six conditions, explicit policy instructions consistently improve separation-of-duties and cumulative-payment judgments, but other effects vary and false-positive rates on matched benign traces remain 22.5–70.4%. The results motivate assigning verification responsibility by evidence need: encode reproducible cross-event rules, connect value claims to authoritative sources, and independently validate remaining LLM-based judgments. This provides a measurement basis for organizational audit and future public-sector assurance requirements.