Order-Marginalized Scoring for Masked Diffusion Models
Arthur Deng ⋅ Sebastian Thrun
Abstract
Masked Diffusion Models (MDMs) are trained under an objective that is symmetric over decoding orders, but at sampling time a single order must be committed to. A common task in such models is to evaluate the log-probability of a candidate sequence, either as a downstream score or to rank a finite pool of candidates. We observe that this evaluation depends substantially on the decoding order chosen, even though the training objective averages over orders. For both a model trained on closed-form Probabilistic Context-Free Grammar (PCFG) data and a pre-trained MDM on OpenWebText, the per-order log-probability rankings of fixed candidate pools disagree on most inputs. We characterize this phenomenon and propose Order-Marginalized Scoring (OMS), a simple Monte-Carlo (MC) estimator of the order-marginalized log-likelihood $\tilde \mu(x)=E_\sigma[\log p_\sigma(x)]$. We use it in two complementary algorithms: a reranker over a fixed candidate pool, and a stochastic best-of-$N$ decoder that combines adaptive position selection with $\tilde\mu$-reranking. On Sudoku, the decoder substantially raises exact-match accuracy over the strongest single-order baseline from $65.3$% to $94.7$%. On a zero-shot protein-stability benchmark, the reranker improves Spearman correlation between log-probability scores and experimental folding stabilities over a per-order baseline by $0.07$. These results suggest that OMS is a useful alternative to per-order scoring for order-agnostic problems.
Chat is not available.
Successful Page Load