Batch-to-Streaming Deep Reinforcement Learning for Continuous Control
Abstract
State-of-the-art deep reinforcement learning (RL) methods have achieved remarkable performance in continuous control tasks, yet their computational complexity is often incompatible with the constraints of resource-limited hardware, due to their reliance on replay buffers, batch updates, and target networks. The emerging paradigm of streaming deep RL addresses this limitation through purely online updates. Existing streaming methods, however, are studied in isolation, trained from scratch until convergence, whereas on-device learning is most useful as a continuation of a pre-trained policy. In this work, we propose two novel streaming deep RL algorithms, Streaming Soft Actor-Critic (S2AC) and Streaming Deterministic Actor-Critic (SDAC), designed as streaming counterparts of SAC and TD3, with performance comparable to state-of-the-art streaming baselines on standard benchmarks. Building on them, we study the batch-to-streaming transition as a problem in its own right: we show that a naive transition can cause a severe and lasting loss of the pre-trained policy, trace the failure to a mismatch between the optimizers used on either side of the switch and show that aligning the optimizers enables streaming deep RL in settings such as Sim2Real.