Fold-Mean Preservation Predicts Weight Merging Success for Cross-Validated Molecular Encoders
SIDDHARTH BETALA ⋅ Luis Pinto ⋅ ALEXANDRE DUVAL
Abstract
Driven by the scarcity of high-fidelity experimental labels, molecular property prediction pipelines routinely produce $K$ fine-tuned checkpoints via $K$-fold cross-validation to maximize data utility. Deploying a single fold wastes most of the trained capacity, yet ensembling all $K$ models is computationally prohibitive at the billion-molecule screening scales now common in drug discovery and materials design. Weight merging promises ensemble-quality predictions at single-checkpoint inference cost, but the toolkit it offers was built for a different problem: constituents fine-tuned on different objectives, whose task vectors conflict and must be reconciled. Folds share an initialization and an objective and differ only in the data they see, so their task vectors are diverse but cooperative, and the behaviour of the existing methods on $K$-fold molecular encoders has not been systematically characterized. Through $5{,}237$ paired evaluations covering 11 merging methods, 5 pretrained encoders spanning three architecture families (SMILES transformers, 2D GNNs, and a 3D equivariant network), and 10 property-prediction tasks under chemically realistic out-of-distribution splits, we find that preservation of the fold-mean update, which is the arithmetic mean of what each fold learned relative to the shared initialization, is the most reliable predictor of merging success (Spearman $\rho = 0.94$, Pearson $r = 0.90$). We formalize this as a composite three-axis preservation score over the direction, coverage, and scale of the merged update, and trace the gain to variance reduction rather than information pooling: merging beats the fold mean on $79.6\%$ of repeats but the best individual fold on only $15.2\%$. Building on these findings, we offer practitioners a deployment recipe based on expected gain and worst-case loss, together with a free pre-flight check, computable from the $K$ checkpoints alone, that predicts whether a given $(\text{model},\,\text{task})$ pair will benefit from merging at all.
Chat is not available.
Successful Page Load