Normalized SGD in the Convex Regime: First High-Probability Guarantees and Momentum Extension
Savelii Chezhegov ⋅ Eduard Gorbunov
Abstract
Normalized stochastic gradient descent (Normalized SGD) is a classical scale-invariant method that updates only along the gradient direction. While Normalized SGD has recently been analyzed in non-convex stochastic settings, its convex theory remains much less developed. We address this gap by studying normalized gradient methods for smooth convex optimization. First, we obtain the theory of deterministic approaches, including first convergence results for momentum variants. Next, we establish first high-probability convergence for mini-batch Normalized SGD with and without momentum, under heavy-tailed stochastic noise with a bounded $\alpha$-central moment. Addressing the key challenge -- bias introduced by normalizing a noisy mini-batch gradient -- we demonstrate that incorporation of a new bias lemma and modified inductive proof allows to control it. As a result, we obtain first high-probability convergence of Normalized SGD in function value, together with explicit oracle complexity bounds.
Chat is not available.
Successful Page Load