Adaptive Behavioral Foundation Models: Continual Training from Reward-Free Experience
Abstract
Behavioral foundation models pretrain general-purpose policies and successor features from reward-free offline data, enabling zero-shot transfer to downstream rewards. However, their zero-shot policies can fail when deployment dynamics differ from those observed during pretraining, and limited online interaction is insufficient to relearn control from scratch. We study reward-free continual training for adaptation to unseen dynamics, assuming that the downstream task is known. We introduce AdaBFM, which frames behavioral foundation model pretraining as learning an adaptive prior over test-time dynamics. AdaBFM combines causal memory with a belief-weighted successor ensemble to condition zero-shot control, and uses task-compatible exploration to collect informative reward-free transitions for continual training. Across unseen PointMaze layouts, continual training improves mean success by 30% on average across baselines relative to their respective zero-shot performance, across 12 locomotion dynamics shifts, it improves normalized return by over 12%. AdaBFM achieves the strongest zero-shot and post-adaptation performance on harder unseen PointMaze tasks, improving by over 17% over the strongest test-time adaptation baseline, and achieves the best or competitive results on harder locomotion shifts.