ProTuon: Progressive Tucker--Muon with Function-Preserving Rank Growth
Abstract
Training neural networks in compressed tensor formats reduces parameter and optimizer-state storage, but fixed ranks constrain model capacity throughout optimization, while naive factor-wise updates ignore the geometry of the Tucker representation. We introduce Riemannian Factors Tucker, which combines geometry-aware Muon updates for the orthogonal factors with Tensorion updates for the higher-order core. Building on this formulation, ProTuon increases mode-wise ranks during training by adding orthogonal directions and zero-padding the core. Each expansion preserves the represented function while exposing new directions for subsequent optimization, allowing training to begin in a compressed form and acquire additional capacity only when needed. Experiments on language-model pretraining show that progressive rank growth outperforms both static Tucker optimization and dense training baselines in final perplexity. Late expansion also recovers near-full-model quality from a compressed checkpoint.