Masked Diffusion Noise Optimisation
Abstract
Masked diffusion language models enable iterative, non-autoregressive generation, but steering frozen models toward sequence-level objectives remains challenging. We introduce Masked Diffusion Noise Optimisation (MDNO), a gradient-based inference-time method that controls an entire denoising trajectory through a single optimisable variable. MDNO parameterises the initial masked noise with continuous, mask-biased logits and optimises them by differentiating a terminal reward through a surrogate of the frozen denoising process. The resulting logits determine the initial state and bias to subsequent denoising steps, providing trajectory-level control without updating either the denoiser or reward model. On open-ended reasoning prompts, we compare MDNO with (compute-matched) Best-of-(N) and FK steering. MDNO achieves the lowest generative perplexity while preserving comparable token entropy, and is preferred to Best-of-(N) by an independent LLM judge. These results support noise optimisation as a viable inference-time steering mechanism for frozen masked diffusion models.