Pessimistic Latent Task-aware Optimization for Robust Offline Meta-Reinforcement Learning
Abstract
Context-based offline meta-reinforcement learning (COMRL) has progressed almost entirely through better task representations: separating task identity from behavior-policy artifacts so that the latent reflects the task information alone. The implicit expectation has been that a clean representation will transfer to out-of-distribution (OOD) tasks. However, on an OOD task the encoder produces a latent outside the training support, where the policy has never been trained. We address this gap with Pessimistic Latent Task-aware Optimization (PLATO), which exposes the policy to off-support latents at training time. PLATO assumes a mild manifold hypothesis: task latents lie on a low-dimensional manifold, with OOD tasks further from the training centroid. It exploits this geometry by perturbing inferred latents outward from the centroid to synthesize counterfactual tasks; a learned decoder ensemble then rolls out short trajectories at the perturbed latent, and the policy is updated on the imagined transitions under a pessimism weight that down-weights steps where the ensemble is unconfident. We prove that the perturbation reaches off-support by a controllable margin, that the epistemic disagreement at the perturbed latent is non-vanishing, and that the OOD generalization gap is bounded by the same perturbation reach. On eight MuJoCo continuous-control benchmarks, PLATO consistently improves out-of-distribution returns over strong COMRL baselines while remaining competitive in-distribution.