AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models
Abstract
Muon trains matrix parameters with a spectral-norm steepest-descent update, but modern token models also contain tied vocabulary tables whose geometry is neither an ordinary dense matrix nor a one-dimensional auxiliary parameter. A tied table receives structurally different gradients from sparse token lookups and dense output classification, and in fully shared encoder--decoder models it may serve encoder input, decoder input, and output roles simultaneously. We propose AF-Muon, an AdamW-free Muon variant that keeps Muon's matrix update, replaces AdamW on tied token tables with a support-aware finite-cap linear minimization oracle, and updates one-dimensional auxiliary parameters with an RMS-normalized direction. Across decoder-only language models, fully shared T5-style models, ImageGPT-style image-token modeling, protein language modeling, and sparse MoE models, AF-Muon improves validation loss over Hybrid Muon and a SCION-style Sign endpoint. Long-horizon and sensitivity studies show that the gains are not a short-horizon artifact or a fragile hyperparameter choice. The results identify tied token tables as a distinct optimizer geometry and show that a simple finite-cap rule gives a robust AdamW-free extension of Muon across models, modalities, and architectures.