An Explicit Basin for Nesterov Acceleration under Local PL near a Manifold of Minimizers
Robert Förster ⋅ Max Obreiter ⋅ Tobias Steinbrecher
Abstract
Over-parameterized models typically have a whole manifold of minimizers, so strong convexity fails and the Polyak–Łojasiewicz (PL) inequality is its natural replacement. On general PL landscapes, however, no first-order method can achieve global acceleration. Hence, the established gain of momentum, $e^{-n\mu/L}$ to $e^{-cn\sqrt{\mu/L}}$, is strictly local. Existing proofs obtain an initialization region from a limit or linearization argument, with no bound on its size, making it non-evaluable in practice. We close this gap by providing an explicit region. Let $\mathcal{S}=\arg\min f$ have reach at least $\tau_0$, and let $f$ be $\mu$-PL with $L$-Lipschitz gradient and $L_H$-Lipschitz Hessian on the tube $\mathcal{S}^{d_0}$, $\kappa=L/\mu$. Then Nesterov's method started at rest, with step $1/L$ and momentum $\frac{1-a}{1+a}$, $a=\sqrt{\mu/(2L)}$, satisfies $\displaystyle \operatorname{dist}(x_0,\mathcal{S})\le\frac{1}{\sqrt{\kappa}}\min\left\lbrace\frac{d_0}{42},\frac{\mu}{84L_H},\frac{\tau_0}{14500}\right\rbrace\ \Longrightarrow\ f(x_n)-f^\star\le 2e^{-n/(3\sqrt{\kappa})}\left(f(x_0)-f^\star\right)$ for all $n$, every iterate remaining in the tube. There is no condition on $\kappa$ and no free parameter. The engine is a Lyapunov function whose tangent frame is re-centred at $\pi(x_n')$ at every step. This makes the affine projection surrogate of Gupta and Wojtowytsch *exact*, and leaves two error terms that classical positive-reach estimates control. Theorem and proof are machine-checked in Lean 4.
Chat is not available.
Successful Page Load