A Million Checkpoints Are Not a Million Samples: Effective Sample Size for Learning from Neural Artifacts
Abstract
Neural-artifact datasets are highly structured: multiple checkpoints belong to the same optimization trajectory, multiple runs may share the same training recipe, and nominal artifact count can therefore substantially overstate statistical diversity. We study effective sample size in a controlled two-domain MLP zoo containing 22,016 artifacts from 1,376 runs and 104 recipes, and replicate the three core tests on public ModelZoo MNIST Hyp-10-rand CNNs. Checkpoint depth exhibits rapid diminishing returns in both populations, and broader independent-run coverage outperforms deeper trajectories at fixed or approximately fixed storage. Under strict recipe-disjoint sampling, local factor classification improves with recipe coverage; public replication supports improved test-accuracy regression with coverage, but not monotone improvement for every target. A constrained hierarchical neural-artifact effective-size index (NA-ESS), evaluated by condition-held-out cross-validation, explains local downstream scaling better than raw artifact count. These results suggest that neural-artifact datasets should optimize and report generating-process diversity rather than checkpoint count alone.