Closure Masks for Hardware-Compatible 2:4 Sparse Muon
Doyoon Kim ⋅ Namhoon Lee
Abstract
Muon is a promising alternative to AdamW, using matrix-valued updates that approximate spectral-norm steepest descent. Can these benefits survive hardware-compatible sparse training? Generic 2:4 sparsity is poorly matched to Muon: its polynomial orthogonalization creates off-support entries, and projecting them away distorts the update geometry that distinguishes Muon from elementwise optimizers. We introduce closure masks, binary masks satisfying the Boolean condition $SS^\top S \le S$, under which Muon’s finite odd-polynomial iteration remains exactly on support. Our construction is transposable 2:4 and alternates partitions across layers to avoid persistent channel separation. Across four GPT-NeoX-style configurations from 31M to 1B parameters, closure reduces validation loss by 5.5–19.4\% relative to balanced random transposable 2:4 under Muon, with no consistent advantage under AdamW. At 160M and 1B, closure-Muon outperforms dense AdamW; at 1B, it remains within 1.7\% of dense Muon. On Llama-3.2-1B, closure improves Muon validation loss by 9.6\%. Diagnostics show zero support leakage and reduced spectral distortion, while isolated Sparse Tensor Core GEMMs achieve $1.29$–$1.52\times$ speedups. Thus, hardware sparsity need not discard Muon’s matrix geometry: the sparse constraint can be designed around the optimizer.
Chat is not available.
Successful Page Load