Online Diffusion Fine-tuning under Doob's $h$-transform Guidance
Zhengyi Guo ⋅ Jiayuan Sheng ⋅ David Yao ⋅ Wenpin Tang
Abstract
Doob's $h$-transform is a probabilistic construction that modifies the law of a Markov process through a harmonic function. In diffusion models, the $h$-function $h(x_t,t)=\mathbb{E}[w(X_0)\mid X_t=x_t]$ alters terminal distributions via the correction term $\nabla \log h$ on the score function. Based on this technique, we propose our \emph{online diffusion $h$-guided fine-tuning} method: we measure an optimality weight for each newly generated sample under the current rollout policy, compute the corresponding $h$-transform terms on the fly, and regress this correction directly into the current flow velocity. We validate the proposed framework across three settings. (1) On a Gaussian-mixture benchmark, online $h$-guidance achieves substantially faster adaptation to a rare target mode than its off-policy counterpart. (2) On contingency-table generation, it more efficiently concentrates samples around prescribed row and column marginals. (3) Finally, on SD3.5-M, online fine-tuning consistently improves reward alignment and adaptation efficiency across multiple text-to-image objectives, compared to baselines such as DiffusionNFT and FlowGRPO.
Chat is not available.
Successful Page Load