How Does Muon Work in Private Deep Learning? Scalable Differentially Private Training via Orthogonalization
Alexander Gaponov ⋅ Grigory Malinovsky ⋅ Egor Shulgin ⋅ Peter Richtarik
Abstract
Training deep neural networks requires optimizers that balance computational efficiency with formal data privacy. While orthogonalized update rules, such as the recently introduced Muon optimizer, have shown strong empirical performance and scalability, their integration with Differential Privacy (DP) remains underexplored. In this paper, we propose convergence analyses of Muon's orthogonalized updates through the lens of a Linear Minimization Oracle (LMO) and establish theoretical convergence guarantees under $(\epsilon, \delta)$-DP constraints. Empirically, we compare DP-Muon with DP-SGD and DP-Adam on language modeling over a wide grid of learning rates, momentum, and batch sizes, and find that the ranking of optimizers changes markedly under privacy: Muon's large-batch advantage in standard training does not carry over to the private setting, where DP-Muon is instead the strongest method at small batch sizes while DP-Adam dominates in the large-batch regime that is optimal for private training.
Chat is not available.
Successful Page Load