Does Muon's Pretraining Advantage Over Adam Survive Fine-Tuning?
Glory Bagai ⋅ Jonas Geiping ⋅ Antonio Orvieto
Abstract
Matrix-aware optimisers such as Muon perform normalised steepest descent in the spectral norm and have been reported to outperform Adam for large-scale language model pretraining. Whether these gains transfer to fine-tuning is less clear: fine-tuning starts from a pretrained solution, uses far smaller datasets, and runs over a much shorter horizon. Additionally, Muon acts only on two-dimensional weight matrices, so embeddings, the language-model head, normalisation scales and biases must be updated by an auxiliary AdamW group with its own learning rate. We study full-parameter fine-tuning of Qwen3-4B under a fixed data budget, sweeping the Muon learning rate, the auxiliary learning rate, weight decay and batch size as separate experimental axes. We find that the initialisation of these hyperparameters accounts for most of the gap: at the default auxiliary rate Muon trails Adam by $0.113$ NLL, but once the auxiliary rate is tuned the two optimisers are indistinguishable ($1.8745$ versus $1.8714$), and this holds across four datasets. Both learning rates grow with batch size at approximately the square-root rate (fitted exponent $0.59$, $r^2 = 0.83$); when only the Muon rate is retuned, performance degrades monotonically with batch size, which may account for part of the reported large-batch weakness of Muon. Decoupled weight decay, by contrast, has no measurable effect (mean paired $\Delta$NLL $= -0.0003$, $95\%$ CI $[-0.0022, +0.0016]$). Our results indicate that, in the fine-tuning regime, Muon matches Adam's performance once both of its hyperparameters are tuned.
Chat is not available.
Successful Page Load