Increasing-Batch AdGD under Glocal Smoothness
Jun-Hyun Kim ⋅ Param Mody ⋅ Curtis Fox ⋅ Mark Schmidt
Abstract
Adaptive Gradient Descent without Descent (AdGD) estimates local curvature from successive gradients, allowing it to take adaptive steps without line-search or a known smoothness constant. Although stochastic variants of AdGD already exist, stochastic noise affects both the update and the curvature estimates, potentially limiting the local adaptivity of AdGD. We study how a growing batch scheme can be used to to control these errors. We analyze stochastic AdGD with an increasing unbiased mini-batch schedule under glocal smoothness, which separates the global smoothness constant $L$ from a smaller local constant $L_\star$. Under strong convexity and glocal smoothness, we explicitly show that progress made during the growing-batch stochastic phase is retained after the batches reach full size, thereby reducing the remaining full-gradient work. Additionally, once the iterates enter the local region, the local phase depends on $L_\star/\mu$ rather than $L/\mu$. Empirically, we compare full-batch AdGD with fixed- and increasing-batch stochastic variants. Across image-classification tasks, slowly increasing batches reach target validation accuracies earlier and attain higher peak accuracy than every fixed-batch baseline. Overall, our results show that growing batches can combine inexpensive stochastic updates with the local adaptivity of AdGD.
Chat is not available.
Successful Page Load