Estimating Rare-Event Probabilities in Masked Diffusion Language Models
Abstract
Estimating the probability that a generative model produces harmful outputs is an important task in AI safety. Yet, estimating these probabilities is particularly challenging when harmful outputs correspond to rare events. While rare-event probability estimation has been studied for autoregressive language models, less attention has been devoted to masked diffusion language models (MDLMs), which attracted growing interest. Existing methods largely exploit the autoregressive factorization of the model distribution, which does not transfer to MDLMs, whose generation proceeds through an iterative denoising process. To address this gap, we introduce a framework for estimating the probability of rare harmful outputs under the original MDLM distribution without modifying the underlying model. The framework uses our new sampler, RollBack sequential Monte Carlo (SMC), to make rare harmful outputs easier to observe and then corrects for this change in sampling when estimating their probability. We compare RollBack SMC with Feynman-Kac Steering, a recent method for steering diffusion-model generation toward high-reward outputs. When used within our estimation framework, RollBack SMC shows lower observed estimation error, higher correction ESS under strong tilts, and lower genealogical concentration, at the cost of additional decoder computation.