Control-centric Representation Learning using Action-free Datasets with Distinct Policies
Abstract
Action- and reward-free data comprises an abundance of real-world in many forms, such as robot manipulation, video game play, and human egocentric recordings. Many works have proposed methods to train feature encoders from these data modalities in order to later fine-tune on less-abundant action-labeled, task-specific data. However, these methods fail when observations contain time-correlated noise. Critically, in many common visual control tasks, a naive encoder cannot distinguish noise from controllable features, leading to poor generalization. In this work, we consider a setting where action-free trajectories are available from agents executing different policies operating in the same environment. We introduce a loss function that provably extracts controllable features and filters out time-correlated noise in the multi-policy video setting. We empirically show its capabilities on distracted versions of control benchmarks like LIBERO, DeepMind Control Suite, and OGBench and we demonstrate capabilities on egocentric human video. This work represents the first scalable method that explicitly extracts endogenous features from high-dimensional, continuous observation and state spaces \textit{without} action labels.