Ethos: Behavior-Aware Selective Distillation for Continual Reinforcement Learning
Yasamin Rajabi ⋅ Narges Kari-DolatAbadi ⋅ Pardis Moradi ⋅ Ali Najar
Abstract
Continual reinforcement learning requires balancing the three central principles of stability, scalability, and plasticity. An agent should preserve previously acquired knowledge, remain capable of adapting to new tasks, and do so under bounded memory and computation. In this work, we formalize a general class of bounded modular continual reinforcement learning methods that maintain and reuse a fixed-capacity pool of task-specific policy components. Building on this formulation, we introduce Ethos, which stores reusable policy parameters, adaptively combines them when learning new tasks, and maintains the bounded pool through behavior-aware policy merging. We compare policies through their action distributions on representative states and merge compatible policies through lightweight policy distillation. We further show that this behavior-aware approach performs better than parameter-space similarity and averaging, leading to more effective merge decisions and improved preservation of previously learned behaviors. On MuJoCo tasks, Ethos improves HalfCheetah $A_N$ by 320\% and reduces forgetting by 58\% over CKA-RL, with consistent gains on Walker2D.
Chat is not available.
Successful Page Load