Optimizer Update Geometry: How Parameter Groups Shape Curvature Exposure
Georg Tirpitz ⋅ Antonio Orvieto
Abstract
We measure loss curvature along the updates of five tuned optimizers during separate pretraining runs of a $162$M-parameter Transformer. At each checkpoint, we construct the next optimizer update on a fixed probe batch. The generalized Gauss--Newton quadratic model expresses the predicted one-step decrease through gradient scale, alignment with the negative gradient, and curvature along the update. We call the last quantity \emph{curvature exposure}. Muon has the lowest exposure, roughly one-tenth of SOAP's. To understand why, we separate the direction within each parameter tensor from that tensor's share of squared update norm. After normalizing by each matrix's mean curvature, Muon and SOAP have the lowest within-matrix exposures. Their much larger difference is in how squared update norm is distributed across tensors. Muon applies its matrix rule to $60$ weight matrices and AdamW to the remaining tensors. These matrices contain $52.4\%$ of the parameters but receive only $5.4\%$ of the squared update norm. They are also sharper on average than the remaining tensors, which produces Muon's negative covariance between squared-update share and diagonal curvature. We derive and verify that Muon's squared matrix-update norm scales with fan-out. The quadratic model and the measured probe-batch loss reduction give nearly identical local rankings, but neither matches final validation loss.
Chat is not available.
Successful Page Load