Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback
Maria Smirnova ⋅ Alexey Kravatskiy
Abstract
SignMuon compresses the Muon update to one bit per parameter by taking the elementwise sign of the Muon Linear Minimization Oracle (LMO) direction: the most direct way to give a matrix-aware optimizer a one-bit budget. It outperforms SignSGD empirically, yet it can ascend even on a linear function. Signing the gradient before the oracle rather than after does not repair this: we exhibit a small explicit instance on which sign-before (MuonUSign) and sign-on-both-sides (MuonSign) ascend as well, so no placement of the sign around the oracle descends in general. Error feedback does not repair sign-after either: applied to the *output* of the oracle it fails for every smoothness constant, step size and momentum. Applied to the *gradient* it does work, and EF21-MuonUSign and EF21-MuonSign attain the standard $\mathcal{O}(T^{-1/2})$ rate for the squared gradient norm on smooth nonconvex problems, the latter at one bit in each direction. Our experiments then reverse this order: across CIFAR-10, its federated split and the nanoGPT speedrun the strongest compressed method is sign-*after*, precisely the placement we prove divergent, with the provably convergent variants behind it.
Chat is not available.
Successful Page Load