Offline Constrained Reinforcement Learning under Partial Data Coverage
Seokmin Ko ⋅ Ambuj Tewari ⋅ Kihyuk Hong
Abstract
We study offline constrained reinforcement learning with general function approximation in discounted constrained Markov decision processes. Existing methods either require full data coverage for evaluating unsupported intermediate policies, are not oracle efficient, or requires the knowledge of data-generating distribution for policy extraction. We propose PDOCRL, an oracle-efficient primal-dual algorithm based on a decomposed linear-programming formulation. The decomposition makes the policy an explicit optimization variable, avoiding policy extraction through the unknown distribution $\mu_D$. We show that naive restricted saddle-point formulations may have spurious saddle points, so realizability of an optimal solution alone is insufficient. We then show that, under a stronger but explicit realizability assumption, every restricted saddle point is optimal, avoiding the regularization and auxiliary function classes used in prior LP-based analyses. These ingredients yield PDOCRL, an oracle-efficient primal-dual algorithm that computes a near saddle point of the empirical decomposed LP Lagrangian and returns a near-optimal, near-feasible policy with a $\widetilde{\mathcal O}(\epsilon^{-2})$ sample guarantee under partial coverage, without access to $\mu_D$ or a reference policy. Empirically, PDOCRL is competitive with strong baselines on standard offline constrained RL benchmarks.
Chat is not available.
Successful Page Load