Results Are Verified — the Harness That Shaped Them Is Not: A Field Study of an Agent-Built Codebase and a Controlled Experiment
Abstract
Teams that build software with coding agents write a harness to keep the agent's work correct: rules, checkers, review agents, acceptance sections. We ask what a project built this way actually verifies. In a 1292-commit, six-month industrial codebase built largely by a coding agent, verification checks the result, not the harness that shaped it: almost nothing in the harness is enforced, and a governance agent built to catch forged test output never ran. This suffices to ship ordinary software, but a project cannot improve a harness it cannot verify. The pass/fail readout that agent evaluations use cannot tell whether a harness changes the work — a gain on the clauses it targets barely moves task success, though a per-clause reading recovers it. On a public benchmark across two model families we wire one acceptance clause four ways — absent, prose, an enforcing checker, and one that only reports success — and the task-level readout cannot tell them apart; we could not show that enforcement beats prose. The differences are real but sit below the outcome: in about a quarter of runs the agent meets the targeted clause while the readout records a failure, and the fake checker passes runs where the clause is unmet.