Separating Real Improvements from Lucky Runs in Autonomous ML Research
Abstract
Autonomous ML research agents increasingly operate by proposing many candidate changes, evaluating them under noisy training, and selecting the strongest observed result. We study how reliably improvements selected by this process survive independent evaluation. In a controlled nanoGPT environment, we first run the same training program across 60 seeds. Selecting the best of 32 runs produces an expected apparent improvement of 0.00544 validation bits-per-byte (valbpb), or 2.7 times our preregistered meaningful-effect threshold delta = 0.002, a gain produced entirely by selection. We then evaluate a frozen 24-candidate bank containing exact-null controls, benchmark interventions, and modifications produced by an autonomous agent without access to the sealed evaluation seeds. Across 749 training runs, nine candidates are exact nulls by construction. Of the remaining 15 candidates, eight independently replicate as meaningful improvements, three fail replication, and four remain inconclusive. The agent finds several large improvements near 0.014 valbpb. In this regime, the effects are large relative to the experimental noise, and preregistered search and confirmation policies make no false declarations on this bank. The harder cases are the smaller improvements near delta, several of which remain unresolved even after 24-33 sealed evaluations. Large effects are readily confirmed, while gains near the scale of experimental noise require substantially more evidence to resolve.