Verifying the Second Agent: Controlled Experiments on Recursive Delegation in Coding Agents
Abstract
A verifier should tell a harness whether adding another agent improves the delivered artifact enough to justify its cost. We test that decision in two within-task crossover campaigns on verifier-scored repository repairs. Recursive trees used roughly six times the tokens at seven nodes and twelve times at fifteen nodes, while no primary verified-correctness contrast resolved a gain over one agent. Execution traces show why added computation need not become complementary work: siblings converged on the same files beyond task locality, and 90.3% of replayable refused sibling patches contained genuine line-level conflicts. We then tested a seven-node single-writer design in which workers returned findings instead of patches. The intervention verifiably removed shared-write integration, but its verified-success difference from shared write was +1.25 percentage points (95% CI [-9.76, +12.82]), leaving no resolved gain at this sample size. Finally, static pre-task routers evaluated out of repository did not outperform the training-chosen constant; one fitted policy selected the single agent on every held-out task. At the tested scales, recursive delegation reliably purchases compute but not a commensurate verified return. Reliable agent development therefore requires verification that joins delivered correctness to resource use and, likely, runtime evidence of whether delegated work remains behaviorally distinct.