Can Muon Adapt Its Stepsize Without Knowing the Solution Distance?
Yury Demidovich ⋅ Abhishek Chakraborty ⋅ Grigory Malinovsky ⋅ Angelia Nedich ⋅ Peter Richtarik
Abstract
Matrix gradient orthogonalization has proved effective in neural-network training, and its trust-region interpretation explains how Muon chooses a normalized direction. It does not determine how far to move along that direction, so a stepsize must still be tuned. Existing theoretical choices may also require the unknown distance $D = \\lVert x_0 - x^\\star \\rVert$ from the initialization to a solution. We develop stochastic Muon methods that choose this directional stepsize without knowing $D$. We formulate directional stepsize selection as a one-dimensional online-learning problem driven by stochastic gradient feedback. For smooth non-convex objectives, a closed-form follow-the-regularized-leader rule gives an expected $O(T^{-1/4})$ stationarity guarantee. For smooth star-convex objectives, a weighted FTRL rule selects the directional stepsizes and, together with a one-gradient-delayed Nesterov direction, yields an expected last-iterate rate of $\\widetilde{O}(T^{-1/3})$. Its parameters use the horizon and smoothness, but neither $D$ nor the oracle variance. On a controlled non-convex matrix problem, the adaptive method eventually outperforms fixed-stepsize Muon baselines retuned for each horizon. On GPT-124M and WikiText-103, practical implementations of both methods achieve validation losses comparable to tuned Muon under matched training conditions.
Chat is not available.
Successful Page Load