Goal-Conditioned RL to Abstract Solved States via Relative Contrastive Reinforcement Learning
Abstract
Goal-conditioned reinforcement learning (GCRL) offers an appealing alternative to tedious reward engineering. GCRL methods uncover the MDP structure of the problem, identifying patterns that would otherwise be difficult to learn from a sparse reward, making them useful for finding complicated paths to known goals, allowing for instance teaching a robot to move an object from the start to goal position. However, in many interesting setups (e.g. Sudoku), both finding the path and knowing the final goal are difficult: if the goal were known, the problem would be solved. Motivated by this, we introduce Relative Contrastive Reinforcement Learning (RCRL), which builds on Contrastive Reinforcement Learning (CRL) to embed trajectories so that the resulting representations move towards their respective final states largely along similar directions. This enables the use of an abstract goal --- which is an artificial goal state that is shared among all the trajectories. We study RCRL on two combinatorial problems, Takuzu and a maze task. We find that the representations learned by a previous CRL-based method for solving combinatorial puzzles lack a shared sense of direction sufficient for solving these problems without knowing the goal at inference time, while RCRL solves both environments (94.1\% on Takuzu, 100\% on the maze).