VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing
Abstract
Inference-time reward guidance is a simple way to improve generative models when completed outputs can be checked or scored by an external verifier. Best-of-N sampling is a particularly strong baseline: it is parallelizable and often highly improves with more samples. However, when the reference model assigns low probability to reward-satisfying samples, selection among independent full rollouts becomes sample-inefficient. We introduce MDM-VGB, a reward-guided discrete diffusion sampler that augments masked generation with value-guided re-masking. MDM-VGB extends the autoregressive VGB backtracking chain from fixed left-to-right prefix trees to any-order masked-state graphs, allowing the sampler to reveal and revise tokens at any position. The resulting Markov chain favors local reveal and re-mask moves that lead to higher-value partial masked states, enabling both root-start reward tilting and leaf-start editing of low-reward samples. We further introduce a flow-cancelled momentum lift of MDM-VGB that preserves the projected target law while reducing oscillatory reveal/re-mask dynamics under finite budgets. We prove that the sampler has the correct reward-tilted leaf law, places non-negligible stationary mass on complete outputs, and remains robust to partial-state verifier error. Across structured reward satisfaction and repair tasks, MDM-VGB improves the quality--cost frontier over reference sampling, best-of-N, and forward-only value-guided rollout.