Verification Without Privilege: Information, Not Architecture, Corrects Multi-Agent Errors
Abstract
When an agentic system generates an important output, the usual approach is to include a review: the agent checks its own work, a second agent reviews it, or the reviewer is provided with rules distilled from past failures. Since the reviewer is seldom treated as a variable in isolation, improvements are often credited to the presence of a reviewer, even if they may originate elsewhere, while null effects remain undetected. We conduct a controlled ablation in a six-node multi-agent pipeline for anti-money-laundering triage, using a 168-case benchmark across three models. Since raw evidence is immutable and every reviewer reconstructs its context from the same fields, the differences in review decisions stem solely from the review logic. An independent reviewer produces no statistically detectable accuracy gain on any model. Injecting failure-derived rules into the same reviewer moves F1 substantially, but on a withheld failure mode, both rule sources fall back to the unassisted baseline. Additional review cycles never recover the single-draft ceiling, and revision consumption saturates well below the permitted cap. Verification helps in proportion to the information the verifier holds that the generator did not, and only over the failures that information covers.