The Oracle Is Mutable: Obedient Destruction of Verification Targets in Shared-Workspace Agent Environments
Abstract
Environment-grounded evaluation gives an agent a writable workspace and verifies what it leaves behind. We document a failure class that needs no adversarial intent. When a benchmark's submission convention places the agent's build output on the pathname of the behavioral reference, following instructions destroys the agent-visible oracle. We study 17 episodes of a program-reconstruction benchmark. Agents overwrote or deleted the reference binary in 6 episodes, under both single- and multi-principal architectures. Three of those episodes then continued differential testing against their own output and logged success: MATCH on every probe, or PASS=39 FAIL=0. The external hidden-test scorer stayed intact throughout. It scored one such artifact at 10.5%. Our claim is deliberately scoped. When an agent-controlled build can replace the reference pathname, agent-side verification can silently become self-comparison, even while the external scorer stays valid. We separate four validity notions that this setting entangles: oracle integrity, agent-side verification, artifact delivery, and external scoring. Predictions recorded before hidden-score access identified every delivery failure (3/3) and no within-band ranking (0/3). We describe a harness-level mitigation: snapshot the reference outside the submission workspace. We state plainly which threat model it does and does not address.