Offline Empowerment Enables Diverse Skill Discovery
Abstract
mpowerment---the mutual information between an agent's choice of skill and its future state---offers a principled formalization of how much control an agent has in its environment, identifying states from which the agent can realize many distinct futures. Previous work has used empowerment as a natural, reward-free objective for discovering a diverse repertoire of skills, as each skill is optimized to realize a distinguishable future outcome. Less attention, however, has been paid to the increasingly common setting in which the agent has access to a dataset of offline, unlabeled trajectories before it begins interacting with the environment online. We propose the first method for estimating empowerment directly from offline data, without any environment interaction, by jointly learning an empowerment-maximizing policy together with its discounted state-occupancy measure. Unlike prior empowerment-based methods, which require on on-policy rollouts from a skill-conditioned policy, our estimator can be trained on trajectories generated from any policy. We demonstrate that our algorithm yields accurate estimates for the empowerment in an environment, unlike prior techniques (when adapted to the online setting). Furthermore, we show that this empowerment estimate can be utilized to derive a maximally empowered policy that provides a diverse repertoire of discrete skills that are useful for downstream task adaptation.