Nowcasting Future Performance for Population Pruning in Survival-of-the-Fittest Training
Abstract
Selecting a neural network from many candidate runs is expensive because each is typically trained to completion, and stopping runs early is risky when learning curves cross and early leaders do not ultimately perform best. We ask whether intermediate checkpoint weights reveal how a model's performance will develop. Across nine image-classification model zoos-CNNs, width-varying ResNet-18s, and Vision Transformers on CIFAR-10, STL-10, and Tiny ImageNet-we characterize ranking dynamics and train performance nowcasters that predict test accuracy five epochs ahead or at the end of training, using observed accuracy and loss histories, engineered weight statistics, and learned SANE checkpoint representations. We then introduce Survival-of-the-Fittest training (SOTF training), which repeatedly ranks active runs by predicted future performance and stops lower-ranked runs, and evaluate retrospectively on held-out complete trajectories whether high performers survive and how much training is avoided. Checkpoint weights contribute signal beyond learning curves, but their utility depends on forecast horizon and population: compact weight statistics are strongest for short-range prediction, structured SANE representations are more effective for final performance, and combining checkpoint and performance features is usually best or close to best. These gains matter most in heterogeneous ResNet-18 and Vision Transformer populations, where rankings change substantially; in homogeneous CNN populations, candidates finish similarly and additional modeling offers little benefit. Where rank changes are consequential, SOTF training preserves high-performing candidates while avoiding a substantial fraction of training-showing that intermediate checkpoints can serve not only as saved artifacts, but as predictive signals for compute-efficient model selection.