When Does Muon Actually Win? A Paired Data-, Time-, and Tuning-Budget Benchmark
Abstract
Muon has reached target loss in fewer updates than AdamW in recent language-model studies, but each update adds matrix orthogonalization. Its value therefore depends on the budget. On a single A100 GPU, we compare a reference single-device Muon implementation with AdamW across 17.7M- and 33.6M-parameter causal language models and 2.70M-parameter vision transformers. A common 23-candidate development budget selects configurations evaluated with five matched seeds, full validation trajectories, isolated optimizer timing, and a frozen one-shot test. Muon reduces validation-loss area under the curve by 3.25--22.44\% at equal data and lowers test loss in all 20 primary pairs. At equal time, Muon improves loss area by 19.68--22.31\% on vision, while AdamW is better by 3.38--6.96\% on language. Muon optimizer steps consume 56.1--69.6\% of training time, compared with 6.1--9.1\% for AdamW, and small tuning budgets penalize Muon more strongly. Within this small-model regime, orthogonalized updates are consistently data-efficient, while the measured wall-clock ordering depends on workload shape, implementation cost, and tuning allowance.