From Curvature to Scaling Laws: Atomic-Norm Regularization Paths in Quadratic Feature Learning
Hee-Sung Kim ⋅ Sungyoon Lee
Abstract
Low-curvature solutions are often associated with good generalization, yet parameter-space curvature alone does not determine which predictor a redundant model selects. In quadratic feature learning the sample-wise Gauss--Newton penalty is exactly the data-weighted trace $\mathrm{Tr}(SC_n)$, a weighted atomic norm on the PSD cone when $C_n\succ0$, and it takes that value at every factorization of the same predictor. A two-factor diagonal model gives a weighted $\ell_1$ penalty up to an explicit balance term. Both statements are finite-sample identities. In the quadratic model the coefficient carries the width: at fixed radius it scales as $p^{-1/2}$ past predictor-class saturation. One curvature budget is therefore not one regularization strength; $\rho\propto p^{1/4}$ holds the induced strength fixed, and training the factorized model separates the two protocols. Read against the replica-symmetric risk predictions of Girardin et al. (2026), that coefficient traces a width-dependent regularization path crossing the source exponent boundary at $p=\Theta(n/d)$.
Chat is not available.
Successful Page Load