Discount-Guided Policy Gradient for Average-Reward Optimization
Ian Hsieh ⋅ Po-An Wang
Abstract
Finite-time analyses of average-reward policy gradient often require a finite global stationary-distribution mismatch coefficient, excluding natural unichain Markov decision processes with transient states. We show that global mismatch control is unnecessary: discounted optimization can guide the policy into a near-optimal region where stationary-distribution mismatch is uniformly bounded. Building on this observation, we propose Discount-Guided Policy Gradient (DGPG), an anytime algorithm that alternates discounted policy-gradient updates with monotone average-reward updates and accepts discounted candidates only when they do not decrease the reported policy's average reward. For finite MDPs in which every stationary policy induces a unichain, we establish a global $O((\log n/n)^{1/5})$ optimality-gap bound and an $O(1/n)$ rate after a finite, instance-dependent localization phase, where $n$ counts all first-order updates. These guarantees hold even when the global mismatch coefficient is infinite. Experiments on two finite unichain MDPs illustrate the benefits of discounted guidance over direct average-reward policy gradient and an $\alpha$-clipped baseline.
Chat is not available.
Successful Page Load