Anytime Convergence via Interpolation Annealing in Schedule-Free Optimization
Abstract
Schedule-Free methods have emerged as strong contenders for eliminating the need to select and tune learning rate schedulers while matching the performance achieved by well-tuned schedules. Recent refinements show that interpolation annealing, in which the Schedule-Free interpolation rate gradually increases, can further improve empirical performance, but whether this mechanism admits convergence guarantees for smooth nonconvex objectives has remained open. In this work, we show that this interpolation annealing mechanism indeed yields an anytime convergence guarantee for Schedule-Free SGD. We also establish an anytime convergence guarantee in the deterministic setting for AMUSE, a recently proposed method that applies the same annealing mechanism to Muon with Schedule-Free averaging. Remarkably, both results hold with eventually constant learning rates and horizon-independent parameter choices.