What Kind of Diffusion Models Do We Need in Online Reinforcement Learning?
Abstract
Score matching and flow matching exhibit powerful expressive capabilities in continuous control tasks. However, since the log probability densities of their generated actions cannot be directly accessed, they introduce substantial challenges to the implementation and optimization of Maximum-Entropy Reinforcement Learning (RL). To address this critical limitation, the RL community has developed a wide range of diffusion-based RL algorithms, leading to a widespread perception that the integration of diffusion models and RL has achieved remarkable success. Yet, is this optimistic conclusion truly well-founded? In this paper, we propose a systematic taxonomy for existing mainstream diffusion RL methods that consists of three distinct categories, and we conduct in-depth theoretical analyses of their inherent limitations and core defects.Our analysis reveals that existing methods either collapse the intended stochastic generative policy into action maximization, require costly long-horizon sampling in high-dimensional action spaces, or rely heavily on proposal distributions whose coverage fundamentally limits policy improvement. Building on the comprehensive analysis above, we propose a novel method \textbf{Copuled Flow} . The core insight of our method is to leverage linear ordinary differential equations to maintain fully tractable log-likelihood calculation, while adopting coupled generation mechanisms to effectively capture complex distributions. In the experiments, we evaluate the \textbf{Copuled Flow} on 3 HumanoidBench tasks, 5 MuJoCo tasks, and 7 tasks from the DeepMind Control (DMC) Suite. Empirically, \textbf{Copuled Flow} achieves superior performance on high-dimensional benchmark tasks compared to competitive strong baselines. Specifically, it significantly outperforms existing diffusion-based RL methods by 100\% on the HumanoidBench benchmark.