Who Verifies the Agents? A Case Series of Unsupported Verification Verdicts
Abstract
We nearly submitted a false example. One agent reported that a released CI tool passed an empty repository. A second installed the package and found that the defect had already been fixed. Our draft described a release we had not tested. We report a case series of six software systems in a coding-agent workflow: five observed verification failures and one source-inspected near-miss. The cases include checks over empty inputs, declarations treated as verified identity, an error page treated as a deployment, and a review queue restricted to its own shortlist. These are related failures of evidence coverage, not six instances of an identical bug. Five systems share an operator and toolchain; the external job-board endpoint is third-party, but its interpretation also belongs to that operator's workflow. The series supports a description of mechanisms, not an incidence estimate or a causal comparison of agent architectures. A further detector encountered three failures while looking for this family of errors, including a failed history read reported as a count of zero. We recommend that verification name its target and coverage, distinguish measurement failure from a negative result, and demonstrate that its control can fail for the intended reason. Independent review helped in these cases, but we do not establish that a second agent is necessary or sufficient. Historical commands and outputs are included; only the external endpoint can be queried by a reader using the materials supplied here. We also describe an unrun comparison design and the limits of its proposed sample size.