Unleashing the Power of Intrinsic-Entropy-Driven Exploration for Off-Policy Generative RL
Abstract
Generative models, such as diffusion and flow-based models, have emerged as powerful policy approximators in deep reinforcement learning (RL) due to their ability to represent multi-modal action distributions. However, their potential for maximum entropy RL remains largely untapped, as most existing methods treat generative processes as black-box samplers or rely on heuristic noise. In this paper, we propose GENIEO, a novel off-policy generative RL framework that leverages the exact intrinsic entropy of an augmented dummy-action policy to drive principled exploration. By designing a structurally invertible affine dual-variable flow, we bypass the numerical approximations of standard ODE solvers. This innovation allows us to rigorously apply the change-of-variables formula to derive a tractable and exact formulation for the augmented dummy-action policy entropy, enabling its direct integration into the off-policy learning objective. Experimental evaluations across 14 challenging tasks from the DMControl and HumanoidBench suites demonstrate that GENIEO achieves state-of-the-art (SOTA) performance. By effectively discovering optimal strategies in complex, high-dimensional environments, our approach bridges the gap between expressive generative modeling and maximum entropy RL.