Scalable Robust-Policy Learning under Total-Variation Uncertainty
Saptarshi Mandal ⋅ Yashaswini Murthy ⋅ R. Srikant
Abstract
Distributionally robust reinforcement learning seeks policies that remain effective under model uncertainty. A recurring computational challenge in robust RL with total variation uncertainty is that robust Bellman updates require global optimization over the state space, which becomes prohibitive in problems with large state spaces and is not automatically alleviated by function approximation. For common-radius state--action rectangular total-variation uncertainty, we show that the minimum-value term in the robust Bellman backup is unnecessary for robust policy learning. Removing it yields a shifted Bellman operator whose fixed point differs from the robust optimal Q-function only by a constant and therefore preserves the optimal greedy policy. We incorporate this shifted update into target-network Q-learning with linear function approximation using a single continuing trajectory, and establish a sample complexity of $\widetilde{\mathcal{O}}(\epsilon^{-2})$, up to approximation errors.
Chat is not available.
Successful Page Load