Beyond the Sampling Frontier: From Evolutionary Discoveries to Generalizable Capability
Abstract
Evolutionary search guided by verifiers can discover solutions to complex problems where standard sampling fails, but its benefits are typically transient and confined to inference-time execution on individual instances. We investigate whether this evolutionary process can be amortized—specifically, whether fine-tuning on solutions discovered by evolution enables cross-problem transfer to unseen hard tasks under direct generation. We first show that evolutionary search provides the greatest advantage on hard problems at the "sampling frontier" (i.e. problems unsolved with naive sampling) finding correct solutions for 79.4\% of frontier problems compared to 51.4\% with compute-matched independent sampling. We then train the base model solely on standalone verified solutions from each method, discarding all search traces and verifiers. On 357 held-out hard problems evaluated with direct generation, the evolution-trained model achieves 1.09\% pass@1 and solves 29 problems at pass@16, compared to 21 for sampling-based self-training and 20 for the base model. These results show that evolution provides transferable training experience, expanding capability beyond standard self-training.