Exact Solutions and Saddle-to-Saddle Dynamics in Nonlinear Matrix Factorization
Abstract
Exact spectral analyses of full training trajectories are well developed for linear networks, but nonlinear activations typically destroy the invariant singular-vector structure that makes them tractable. We identify a nonlinear matrix-factorization setting in which this structure survives. For targets with orthogonal columns and elementwise activations, spectral initialization confines gradient flow to an invariant diagonal manifold, reducing training to independent scalar ODEs that are exactly solvable by quadrature. For locally linear activations, small balanced initialization yields sequential mode emergence, with larger target singular values learned earlier, producing trajectories near a sequence of increasing-rank critical points. Small random initialization exhibits qualitatively similar singular-value dynamics in this model. These results provide an exactly tractable nonlinear setting for studying stage-like feature learning and saddle-to-saddle dynamics.