The Verifier Never Saw the Error: Canary Audits for Document Agents
Abstract
Many document agents read a page once, convert it to structured text, and verify only that text. This architecture has a blind spot: the reader can remove an error before the verifier sees it. We expose the blind spot with canary bank statements. Each statement contains one printed balance that deliberately disagrees with the surrounding transactions. Copying the balance preserves the evidence and makes an arithmetic checker fire; replacing it with the value implied by the ledger makes the same checker accept. On a fresh 36-document replication, Opus 5 produces this fully silent failure on 10 documents (27.8%, exact 95% interval 14.2 to 45.2) under a neutral extraction prompt and 3 under an explicit verbatim prompt. On a matched 48-document set, the same 36 fresh pages plus the 12 discovery pages with one output per prompt, GPT-5.6 fails on 43 (89.6%, 77.3 to 96.5) and 6 documents, respectively. The strongest experiment holds the target-cell pixels fixed while changing only the surrounding ledger. Opus 5 reports the context-implied value in all 12 neutral-prompt variants, even though the visible target is unchanged; under the verbatim prompt, only 1 of 12 answers follows the ledger. An OCR-specialized model shows no context-repair signature in 71 parsed runs. These are controlled audit rates, not deployment prevalence. They show that a report-only verifier can be correct about its input yet wrong about the page. We conclude with a concrete interface design: retain the source crop, literal transcription, inferred value, and uncertainty as separate evidence.