The Local Geometry of Magnitude and Direction Decoupling
Abstract
Recent advances in optimizer design decompose the matrix-valued weights into magnitude and direction components, accelerating the training of Transformers. While the empirical evidence points to a practical advantage, the theoretical understanding of this acceleration is unclear. In this work, we start by studying the effect of gradient decoupling using a simple quadratic model. For this convex case, For strictly convex quadratics, the optimal two-parameter decoupled step achieves at least as much one-step decrease as gradient descent with exact line search. Next, we train a GPT model of 160M parameters on 3.2B FineWebEdu tokens, and use line search to compare the loss along different directions. In this non-convex setting, we empirically observe that line search along the decoupled direction matches or improves on line search along the gradient direction. Our results offer a optimizer-agnostic understanding of the decoupling advantage.