Landscape of Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action
Xin Chen ⋅ Minda Zhao
Abstract
We study the optimization landscape for infinite-horizon discounted Markov decision processes (MDPs) with general state and action spaces under structured policy classes. Existing approaches to establishing global convergence guarantees for policy gradient methods require the parameterized policy class to be closed under weighted policy improvement for every policy in the class, a property that may fail even when the class contains an optimal policy. To address this issue, we propose weaker conditions that guarantee the absence of suboptimal stationary points and establish the Polyak--Lojasiewicz--Kurdyka (PLK) condition for policy optimization, which guarantees the global convergence of first-order methods. Our general results apply to settings already covered by existing analyses, including tabular MDPs and linear quadratic regulator problems. We further verify our proposed conditions for two operations models outside the scope of existing analyses: inventory systems with Markov-modulated demand and stochastic cash-balance problems. For both models, we establish exponent-one and, under additional curvature assumptions, exponent-two PLK conditions, which, under smoothness, yield an $\mathcal{O}(1/\epsilon)$ iteration complexity and linear convergence, respectively, for projected gradient descent using exact policy gradients. To the best of our knowledge, we provide the first non-asymptotic convergence rates for solving infinite-horizon discounted inventory systems with Markov-modulated demand and stochastic cash-balance problems using policy gradient methods.
Chat is not available.
Successful Page Load