Orthogonal Updates for the Win: Towards Accelerated Adaptive Minimax Optimization
Zhiwei Zhai ⋅ Xinyu Wang ⋅ Wenjing Yan ⋅ Lei Ding ⋅ Ying-Jun Zhang
Abstract
Single-loop minimax methods are appealing for modern machine learning, but standard stochastic gradient updates can be unstable and oscillatory under tightly coupled primal--dual dynamics, making performance notoriously sensitive to stepsizes. To fundamentally improve stability, we develop an orthogonal update framework tailored to minimax optimization, which applies updates along orthogonalized directions to reduce update anisotropy. However, orthogonalization reshapes the optimization geometry and makes coordinating primal--dual stepsizes even more delicate. To overcome this challenge, we develop \textbf{AdaSGDA}, which equips orthogonal updates with an adaptive stepsize rule driven by accumulated gradient norms. This mechanism automatically balances coupled primal--dual progress under orthogonalization, yielding an effectively parameter-free method for nonconvex-strongly-concave minimax optimization. We prove that AdaSGDA avoids problem-dependent tuning while attaining the state-of-the-art $\mathcal{O}\left(T^{-1/4}\right)$ convergence rate. We further propose \textbf{AdaSGDA-VR}, which incorporates a carefully designed variance reduction scheme and achieves the faster $\mathcal{O}\left(T^{-1/3}\right)$ rate. Extensive experiments demonstrate that our methods deliver competitive performance compared to existing minimax optimization methods.
Chat is not available.
Successful Page Load