Monarch–Muon: Structure-Aware Muon for Fast and Memory-Efficient Pretraining
Abstract
Matrix-aware optimizers such as Muon improve language-model pretraining, but they orthogonalize dense momentum matrices, so both the optimizer state and the optimizer step scale with the full weight. We introduce Monarch--Muon, which moves the matrix-aware update onto a Monarch-structured parameterization. We show that the Muon trust-region subproblem separates exactly over block-diagonal factors, so the optimizer never forms or orthogonalizes a dense matrix and instead runs independent Newton--Schulz iterations on the trainable Monarch blocks. The reduction comes from structured connectivity rather than a rank constraint, so the represented weights remain full-rank. We prove nonconvex convergence for the true factor gradient under explicit smoothness and stochastic-gradient assumptions, including a finite-step Newton--Schulz alignment bound that needs no positive lower bound on the nonzero singular values. The block count sets the trade-off directly: Monarch-Muon remove 10\% of the trainable parameters for a 2.3\% increase in validation loss over dense Muon, and four blocks remove 37\% for 4.7\%. Optimizer state memory falls by 4.1x against dense AdamW and 2.2x against dense Muon at 6.89B, and the margin widens monotonically with hidden dimension, from 1.9x at 257M to 4.1x at 6.89B, so Monarch--Muon holds the lowest optimizer state of every method we measure at all five scales. Step time at 6.89B drops from 2986.9 ms to 1638.6 ms against dense Muon, and under TP=2 tensor parallelism the optimizer phase runs 3.3x faster than duplicated dense Muon at 34.93 GB peak memory against 54.45 GB.