Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity
Shira Vansover-Hager ⋅ Matan Schliserman ⋅ Ofir Schlisselberg ⋅ Tomer Koren
Abstract
Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training. This raises a basic robustness question, crucial to reproducibility and reliability: how sensitively do MD dynamics depend on their inputs? We focus on initialization, often itself a pretrained or previously aligned model. Quadratic-regularized MD, including GD and Mahalanobis geometries, is well-known to be stable for convex smooth objectives. We show a sharp contrast: once the regularizer is non-quadratic, MD can be exponentially more sensitive to initialization than GD, even with a well-conditioned regularizer in Euclidean norm. We give a three-dimensional construction with a convex, smooth objective and a strongly convex, smooth, well-conditioned regularizer where an initial $\varepsilon$ perturbation is quickly amplified to $ \min\\{ \text{polylog}^{-1}(1/\varepsilon), \varepsilon e^{\Omega(\eta T)} \\} $ after $T$ iterations of MD with step size $\eta$. For canonical KL-regularized MD on the simplex, we show that even linear objectives can amplify an initial $\varepsilon$ perturbation exponentially fast in high-dimensional or near-boundary regimes. Finally, we propose Anchored MD, which adds a Bregman term to a fixed point, and show it achieves $O(1/\sqrt{T})$ stability, while preserving optimization guarantees up to logarithmic factors.
Chat is not available.
Successful Page Load