A Radial Smoothness Model for LMO-Based Optimization
Sam Laing ⋅ Egor Shulgin ⋅ Antonio Orvieto ⋅ Peter Richtarik
Abstract
Recent optimization algorithms such as Muon and sign gradient descent select structured update directions through Linear Minimization Oracles (LMOs) over norm balls, while their update radii are typically specified through external learning-rate schedules. Motivated by this separation, we study radial upper models in which the first-order Taylor remainder is controlled by a function $\phi(\|\Delta x\|)$. Minimizing such a model decomposes into an LMO direction and a one-dimensional optimization problem for the radius, governed by the dual norm of the gradient. Under mild convexity and growth conditions on $\phi$, we establish a descent inequality and a nonconvex first-order stationarity guarantee, and give an explicit specialization to power-law models. We then use the same radial construction more broadly as an optimizer-design principle under operator-norm geometry, comparing several radius rules on a neural-network testbed and a $30$M-parameter transformer. The experiments show that the radial rule can substantially affect optimization dynamics, and that sublinear dependence of the radius on the dual-gradient magnitude can be competitive with a tuned fixed-radius LMO baseline.
Chat is not available.
Successful Page Load