Meta-Reinforcement Learning with Zero-Shot Reinforcement Learning
Abstract
Meta-reinforcement learning (meta-RL) agents adapt to novel tasks from test-time experience, but require diverse training sets of environments and reward functions that are expensive to construct. Behavior Foundation Models (BFMs) learn policies from reward-free data that are capable of zero-shot RL (ZSRL), adapting to new objectives at test time when given a large reward-labeled dataset. Recent work has shown that BFM task inference can be performed online, making BFMs and meta-RL direct competitors for generalization in fixed environments. We ask whether BFMs can be extended to adapt to novel reward functions in novel environments, and identify four key limitations: environment identification, online data collection, test-time exploration-exploitation, and mixed reward supervision. We study these subproblems through toy domains, standard BFM benchmarks, and simulated humanoid locomotion, then propose a unified framework that addresses all four. The resulting method, MetaBFM, is a hybrid RL agent spanning meta-RL, ZSRL, and intrinsic exploration. During training, MetaBFM combines supervised reward-following with unsupervised reward-free learning; at test time, it interpolates between exploration, exploitation, and ZSRL-style reward inference. We evaluate MetaBFM on two toy meta-RL domains and at scale in MetaWorld, showing that hybrid meta-RL/ZSRL agents can learn more general behavior from the same set of reward functions and may reduce the need for future meta-RL domains to hand-design diverse training sets.