Does U-W Ratio Alignment Benefit Muon in LLM Pretraining?
Abstract
Most of the top-performing entries in the modded-nanogpt optimization benchmark---a leading competitive benchmark for efficient optimizer design in LLM pretraining---now use a "u/w floor" rather than decoupled weight decay. The u/w floor imposes a common lower bound on the ratio of the Frobenius norm of a parameter block's update matrix to that of its weight matrix (u-w ratio). The realized u-w ratios concentrate at this bound for most training steps, inducing u-w ratio alignment across blocks. To determine whether this alignment contributes to the u/w floor's empirical advantage, we study MuonL and typed MuonL (t-MuonL), which isolate the alignment effect by rescaling Muon updates while preserving their directions. MuonL assigns a common ratio across all Muon-updated blocks, whereas t-MuonL aligns ratios only among blocks sharing the same Transformer role. For both variants and original Muon, we derive a measurable diagnostic based on the one-step loss decrease and establish covariance-based sufficient conditions under which alignment is beneficial. When the aligned matrix blocks share the same shape, the condition favors alignment when blocks with larger weight norms tend to exhibit greater first-order descent along Muon directions and lower adjusted directional curvatures, provided these gains outweigh the rescaling-factor variance. For alignment across differently shaped blocks, the criterion additionally accounts for matrix dimensions. Diagnostics along a Muon trajectory from the benchmark favor both MuonL and t-MuonL, consistent with independently trained runs in which both achieve lower validation loss than Muon. These results support u-w ratio alignment as a mechanism behind the u/w floor's empirical advantage and a useful intervention for Muon-based LLM pretraining.