Momentum Smooths the Path to Gradient Equilibrium
Abstract
Online gradient descent has recently been shown to satisfy gradient equilibrium for a broad class of loss functions. This means that the average of the gradients of the losses along the sequence of estimates converges to zero, a property that allows for quantile calibration and debiasing of predictions, among other useful statistical properties. A shortcoming of online gradient descent when optimized for gradient equilibrium is that the sequence of estimates is jagged, leading to volatile paths. In this work, we propose the generalized momentum method, defined as a weighting of past gradients, as a broader algorithmic framework with guarantees to smoothly postprocess (e.g., calibrate or debias) predictions from black-box algorithms, yielding estimates that are more meaningful in practice. We prove it achieves gradient equilibrium at the same convergence rates and under assumptions similar to those of plain online gradient descent, all the while producing smoother paths that preserve the original signal amplitude. Of particular importance are the consequences for sequential decision-making, where more stable paths translate to less variability in statistical applications. These theoretical insights are corroborated by experiments on real data, showcasing the benefits of adding momentum.