Muon Moves the Boundary of Transformer Trainability
Abstract
Muon has moved from small-model studies into trillion-parameter training, but that transition required a new control for attention-logit instability. We study how Muon’s spectral map changes the boundary between runs that learn and runs that do not. Our control is Tp(M ) = U ΣpV ⊤ at fixed root-mean-square (RMS) update magnitude: p = 0 is the polar map, p = 1 is raw momentum, and intermediate values partially flatten the spectrum. Locally, flattening changes the update con- dition number from κ to κp while increasing weak-direction sensitivity as κ1−p . Finite Newton–Schulz depth limits this sensitivity below a cutoff proportional to 3.4445−k ; the measured cutoff slope agrees within 0.2%. In a tanh-network toy, p changes the box-counting dimension of the trainability boundary from 1.86 at p = 0 to 1.37 at p = 1. One Newton–Schulz step comes within 0.02 of the exact- polar dimension, and the measured dimension grows with training horizon for the four plotted controls. In Transformers, the same exponent changes model-size transfer. Rates calibrated on a 10M-parameter model are reused without retuning on a 30M model; p = 0.75 learns on all three tested seeds, while AdamW and p ∈ {0, 0.5, 1} diverge on all three. Near the 10M Muon boundary, 20 of 25 seed-aligned cells mix learning and nonlearning outcomes, and 103 of 200 labels change between 128 and 256 updates. Thus the toy supports measured noninteger boundary geometry, whereas the Transformer evidence supports a probabilistic, horizon-dependent boundary rather than a fractal-dimension claim.