NanoFold: Designing Reproducible Protein Structure Benchmarks through Principled Sampling
Chris Hayduk ⋅ Krithik Ramesh
Abstract
Protein structure prediction has progressed rapidly, with a flourishing ecosystem of open-source AlphaFold (AF)-style systems delivering real progress. But isolating which architectural and training choices actually drive that progress is difficult. Production-scale training is both computationally prohibitive and not accessibly reproducible; moreover, the public corpora are complex, with structural biases that compound in parameter- and compute-constrained regimes. To address these gaps, we introduce NanoFold, a compact, fixed-data benchmark for AF-style training studies. Its sealed-hidden tracks address three fundamental questions, with a $\textit{limited}$ track for sample efficiency, a $\textit{research large}$ track for whether early gains persist under more optimization, and an $\textit{unlimited}$ track for best fixed-data final performance. NanoFold's split is designed rather than naively sampled, with chains grouped into biological units under MMseqs2 cluster and PDB-entry disjointness, stratified across structural metadata, and allocated to $10{,}000$ training, $1{,}000$ public-validation, and $1{,}000$ sealed hidden chains. We verify the construction through statistically principled diagnostics, including a randomization study over $1{,}000$ alternative valid splits, confirming the dataset is diverse, well-distributed, and statistically typical of its constraint class. We use NanoFold to run comprehensive experiments on model behavior across scales and training regimes, demonstrating that the benchmark is learnable but unsaturated, scales predictably with budget, cleanly separates training primitives, and that the underlying codebase flexibly supports head-to-head comparison. NanoFold makes architectural and training progress in protein structure prediction more transparent, comparable, and reproducible at a compute-accessible scale.
Chat is not available.
Successful Page Load