Towards Convergence of PPO: An Approximate Descent Approach
Abstract
Proximal Policy Optimization (PPO) and it's variants are the most widely used policy-gradient methods in reinforcement learning, yet the role of its actor update mechanism is still poorly understood theoretically. In particular, standard PPO combines clipped surrogate gradients, multiple epochs of minibatch updates, and reuse of rollout data, but existing convergence analyses do not fully capture this update structure. In this work, we theoretically study PPO policy updates with symmetric clipping from a policy-gradient perspective and interpret its actor updates as a cyclic biased-gradient method with sample reuse and random reshuffling. Our first contribution is a clean formalization of PPO clipped actor updates through surrogate gradients that approximate the true policy gradient. Using the performance difference lemma, we prove a linear bias bound which quantifies how the surrogate gradient drifts as the policy moves away from the sampling policy. Our second contribution is a convergence analysis of cyclic surrogate-gradient ascent, showing that additional PPO-style biased updates can improve progress under conservative learning rates without requiring additional samples. Finally, we analyze the stochastic minibatch version with reshuffling and obtain convergence-to-stationarity guarantees under standard smoothness assumptions and a bounded critic-bias condition. Overall, our results provide a theoretical interpretation of PPO’s multi-epoch actor updates: extra clipped surrogate steps introduce bias, but can still improve optimization efficiency by compensating for small, stable step sizes through sample reuse.