A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models
Abstract
We formulate reinforcement learning (RL) in continuous time over discrete state spaces via a stochastic control approach, in which the state dynamics form a controlled continuous-time Markov chain, and derive the corresponding policy gradient, yielding continuous-time PPO and GRPO. As the primary application, we obtain a continuous-time RL (CTRL) framework for fine-tuning score-based discrete diffusion models by treating the discrete score as the action. It needs no differentiable reward and, unlike existing GRPO variants using terminal rewards only, admits intermediate rewards along the whole denoising trajectory. Specialized to masked diffusion models, it yields policy parameterizations over the vocabulary simplex with analytically tractable probability ratios. We showcase the effectiveness of our methods on RL post-training of LLaDA-8B-Instruct on mathematical reasoning and coding tasks.