Discovering Compositional Primitives Invariant to Context Shift
Hochan Bang ⋅ Ganghun Lee ⋅ Min Whoo Lee ⋅ Minsu Lee ⋅ Byoung-Tak Zhang
Abstract
Unsupervised reinforcement learning aims to acquire diverse and reusable primitives without task-specific rewards. However, existing methods often condition skill policies on the full state of the environment to support exploration, which can make the learned skills coupled to the global context and thus sensitive to context changes. Moreover, executing a single episode-level skill can involve multiple primitive behaviors, but explicit control over those primitives is often not available. In this work, we propose a framework for discovering primitives with improved reusability, and provide proof-of-concept evidence from experiments with MuJoCo $\text{Ant}$ in different geometries. Our framework decomposes a skill-conditioned policy into a global module conditioned on the full state of the environment, and a local module conditioned on the egocentric observations. Empirical results suggest that our framework yields primitives that are invariant to context shift while preserving exploration coverage and improving downstream task performance.
Chat is not available.
Successful Page Load