The Memory of Failure Should Be Lossy: Wrongful Convictions in Shared Failure Banks for Autonomous Research
Abstract
Autonomous research systems increasingly share negative knowledge: failed attempts enter a common bank so no agent repeats them. We show this design, in its standard configuration—irreversible binary negative verdicts, re-exercised positives, success-gated proposal—has a structural flaw. Failure verdicts are statistical claims from noisy, small-n experiments, and under that configuration their error types are asymmetric: a false confirmation is built upon and self-exposes, while a false conviction is locked precisely so nobody retries it. In a model of cumulative research we prove that with exogenous supply and binary records full deference is optimal, but with success-gated supply discovery is inverted-U in deference, with losses up to 81% at long horizons. Measured wrongful-conviction rates track the closed form across three landscapes—10--20% in live training loops, 35--50% for verification-stage near misses (the most evidence-laden negative verdicts are the least reliable)—and a 122,817-conviction public benchmark confirms that marginal mass, not aggregate noise, sets the rate. Behaviorally, the lock replicates across twelve deployed LLMs from nine families and two domains: the DEAD label alone suppresses selection of an a-priori-best idea in every family (free-choice revisits 0/94), beliefs dissociate from choices, and fair prompt-level repairs are either family-unreliable or indiscriminate—the record, not the prompt, is the calibrated intervention point. Inside gated loops, code-enforcing the lock strips binary's soft self-repair (rescues 14→0) while graded's open records survive it; in a controlled end-to-end loop with an algorithmic proposer, evidence-graded memory delivers 30--32% more truth-validated discoveries than binary locks, no memory, or volume-matched parole, ceding its edge in a low-noise cell exactly where the phase structure predicts. The fix is not less sharing but calibrated conviction: evidence-threshold locking—the Bayes rule for the induced one-step locking decision—is a schema change, not an architecture change.