Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
Bingbin Liu ⋅ Rachit Bansal ⋅ Depen Morwani ⋅ Nikhil Vyas ⋅ David Alvarez-Melis ⋅ Sham Kakade
Abstract
Diagonal preconditioners are computationally tractable approximations to second-order optimizers and have been central to the recent push for faster training of deep models. Two predominant families are based on Adam and on the Gauss-Newton (GN) matrix: Adam tracks running statistics of gradients, while GN-based methods (e.g. Sophia) use the diagonal of the Gauss-Newton matrix. We compare these two families through the lens of two factors: the basis in which the diagonal preconditioner acts and the impact of gradient noise from mini batching. Using linear and logistic regression as analytic testbeds, we show that (i) regardless of basis, there exist instances where Adam outperforms both $GN^{-1}$ and $GN^{-1/2}$ in the full-batch regime---even in the GN eigenbasis for logistic regression---and (ii) in the stochastic regime, Adam behaves equivalently to $GN^{-1/2}$ for linear regression under Gaussian inputs. These predictions are corroborated by experiments on convex and non-convex problems including MLPs, CIFAR-10, and Transformers.
Chat is not available.
Successful Page Load