When to use what Schatten-$p$ norm in deep learning?
Thomas Pethick
Abstract
Schatten-$\infty$ based optimizers such as Muon have shown promising empirical performance, but there remains seemingly conflicting observations regarding whether they are beneficial. We resolve this conflict by showing that the conclusion is regime dependent. Even when the problem geometry is naturally the Schatten-$\infty$ norm, smaller Schatten-$p$ geometries can be optimal, specifically in the low-dimensional regime, which we show includes Chinchilla scaling. This conclusion follows from a new noise-robust acceleration result for the SODA framework for $p>2$.
Chat is not available.
Successful Page Load