When Acceleration Breaks: (Stochastic) Edges of Stability in Normalized SGD
Itay Lavie ⋅ Clarissa Lauditi ⋅ Cengiz Pehlevan
Abstract
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, ranging from coordinate-wise adaptive methods to matrix-normalized optimizers such as Muon. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. The fixed-norm updates induce an effective learning rate that grows as the gradients shrink. We develop a dynamical mean-field theory (DMFT) that describes the joint dependence of the loss on training time, model width, and batch size. Normalization initially accelerates SGD: when SGD achieves power-law convergence with exponent $r_{\rm SGD}$, normalization transforms the temporal exponent into $2r_{\mathrm{SGD}}/(1-r_{\mathrm{SGD}})$ for $r_{\mathrm{SGD}}<1$, while for $r_{\mathrm{SGD}}=1$ and $r_{\mathrm{SGD}}>1$ it yield exponential and formal finite-time convergence, respectively. At finite step size, however, the growing effective learning rate makes the acceleration self-limiting and ultimately drives the dynamics to marginal stability. The late-time DMFT yields a width--batch phase diagram with three regimes: a finite-width bottleneck produces an irreducible error, while in the high-dimensional limit increasing batch size suppresses sampling fluctuations and drives a phase transition from a noise-sustained edge of stochastic stability to a deterministic edge-of-stability floor. Thus, model capacity, minibatch noise, and discrete step stability threshold emerge as competing late-time bottlenecks within a single solvable model.
Chat is not available.
Successful Page Load