Muon Under Covariate Shift: Spectral Learning and Temporal Stability
Abstract
Muon's advantage over AdamW and momentum SGD can grow under covariate shift, even when their in-distribution errors are similar. We investigate this behavior through controlled regression experiments and an aligned matrix model, which shows how spectral updates learn directions that source-weighted gradients underemphasize. Because generalization also depends on sensitivity to training data, we complement this account with an analysis of neighboring training trajectories. Across controlled nonlinear tasks and real tabular datasets, Muon produces smaller accumulated prediction perturbations, yet a larger fraction of these perturbations persists until training ends. This balance usually favors Muon, but synthetic counterexamples show that persistence can outweigh the reduction in magnitude and reverse the stability ordering. Understanding optimizer-dependent generalization therefore requires examining both the directions learned and how perturbations accumulate over training.