Can Acceleration in Gradient-Norm Minimization Be Anytime? Sharp Last-Iterate Limits in Smooth Convex Optimization
Pierre Vernimmen ⋅ François Glineur
Abstract
In smooth convex optimization, the gradient norm is a directly observable measure of stationarity. Accelerating a first-order method that minimizes the gradient norm is known to be more delicate than accelerating the minimization of function values. Optimal accelerated methods such as OGM-G (Kim & Fessler, 2021) exist for any prescribed finite horizon, but their coefficients depend explicitly on the length of that horizon, i.e. the number of iterations. We ask what kind of acceleration remains possible when the stopping horizon is unknown to the method, i.e. for horizon-independent methods. Diakonikolas & Wang (2022) conjectured that an $\Omega(N^{-1})$ lower bound on the squared gradient norm holds at every horizon $N$ for any nonadaptive, horizon-independent linear-span first-order method, and Tsai et al. (2026) proved it for gradient descent with positive stepsizes. We show that the conjecture is false in the broader linear-span class of methods: a single horizon-independent method can achieve near-$N^{-2}$ last-iterate guarantees on a density-one set of horizons. We show instead that an $\Omega(N^{-1})$ lower bound must hold for infinitely many horizons. More precisely, if $\mathcal{G_\mathnormal{N}}(\mathcal{A})$ denotes the squared-gradient-norm guarantee after $N$ iterations for a method $\mathcal{A}$, we prove that $\limsup_{N\to\infty}N\mathcal{G}_N(\mathcal{A}) \ge 1/2$ for any method $\mathcal{A}$. This bound is sharp: the constant $1/2$ is exactly attained by the horizon-independent gradient-descent schedule of Rotaru et al. (2026). In addition, we show that the above two extreme behaviors cannot be achieved by the same method: any method $\mathcal{A}$ with an $o(N^{-1})$ guarantee on a subsequence of iterates must satisfy $\limsup_N N\mathcal{G}_N(\mathcal{A})=\infty$. Thus, horizon independence allows acceleration at almost every horizon, but not uniformly accelerated last iterates. In contrast, best-so-far output admits a uniform $O(N^{-2})$ guarantee.
Chat is not available.
Successful Page Load