When Does Pruning Earn Its Optimization Cost? A Selection-Valid PAC–Bayes Diagnostic for Frontier Optimizers
Ahanaf Ariq
Abstract
Can an optimizer beat Adam once dynamic pruning and statistical selection costs are counted together? We study this question without claiming a universal winner. We introduce a finite-family, split-prior diagnostic that treats an optimizer, pruning history, curvature geometry, and posterior scale as one predeclared candidate. The candidate's mean and nested support path are fixed on $S_{0}$; an independent slice $S_{1}$ estimates a local curvature matrix, evaluates a bounded posterior risk, and selects after charging the complete candidate prior mass. The resulting PAC–Bayes numerator separates trajectory description length from Gaussian curvature, family selection, and confidence. We derive the exact KL for a full-rank Fisher-shaped Gaussian and its soft effective dimension. On three real classification datasets and four compact NumPy optimizer baselines, the charged rule is selection-valid but usually clips at one: it frequently selects a dense Adam candidate because gradual pruning costs roughly 3,168 nats to describe. In contrast, selecting by $S_{0}$ loss area chooses a Newton–Schulz momentum baseline on two datasets and achieves higher held-out accuracy in this small-model study. The result is a reproducible frontier diagnostic: an optimizer can win the optimization frontier while losing the certificate frontier, and neither frontier should be presented as the other.
Chat is not available.
Successful Page Load