From Mixed Futures to Closed-Loop Cycles: Prediction and Policy Extraction in Offline Goal-Conditioned Reinforcement Learning
Abstract
Offline goal-conditioned reinforcement learning uses recorded trajectories to learn how to reach a goal. However, predictions about those trajectories may not describe what happens when a new policy chooses its own actions. We investigate this mismatch in contrastive reinforcement learning using a controlled maze. Actions appear favourable because their recorded continuations eventually reach the goal, but repeatedly selecting the highest-scoring action creates a cycle. Changing how the same data are sampled can remove the cycle. Training a policy to select actions can also produce successful behaviour where direct selection from its critic's scores fails, although behavioural cloning alone achieves the same success rate. Under partial observations, an exact Bellman calculation also becomes vulnerable to cycling. These findings show why evaluating a critic's predictions and evaluating the policy that uses them are distinct requirements for understanding offline RL behaviour.