Small Batch Noise Restores Simplicity Bias of Muon
Abstract
Simplicity bias, the tendency of training dynamics to learn features sequentially, has been proposed as a general explanation for the success of deep learning. An important driver of this phenomenon is saddle-to-saddle behavior, in which the gradient flow trajectory remains close to a sequence of saddle manifolds while learning one feature at a time. However, Muon, a state-of-the-art optimizer based on spectral descent, does not exhibit such dynamics in the full-batch setting and instead learns features in parallel. How is it then that Muon is so successful in deep learning settings? We show that stochastic gradient noise fundamentally changes this behavior: while full-batch Muon lacks simplicity bias, small-batch Muon recovers it and exhibits sequential feature learning. We explain this transition through stochastic smoothing of the spectral descent operator underlying Muon. When gradients are small relative to the noise scale, the expected spectral update becomes approximately proportional to the actual gradient, which restores gradient-flow-like dynamics. We formalize this mechanism for deep linear models, by deriving the effective continuous dynamics for noisy spectral descent. Experiments on low-rank matrix reconstruction and vision transformers support these findings.