Rethinking Diffusion Decoding via Structural Commitment
Abstract
Standard decoding in diffusion language models (DLMs) is controlled through parallel, token-wise unmasking decisions. However, diffusion decoding is inherently a structured, sequence-level process, in which token states evolve under dynamic cross-token dependencies and are therefore poorly captured by purely token-wise unmasking control. Motivated by this mismatch, we view diffusion decoding from a structural perspective, in which iterative denoising progressively gives rise to token subsets that are internally reliable and weakly dependent on the remaining masked positions. We therefore introduce structural commitment, a sequence-level unmasking principle that treats approximately closed token subsets as the basic unmasking units in diffusion decoding. We instantiate this principle with Structural Commitment via Closure Expansion (SCCE), a structure-aware inference-time algorithm that discovers approximate closures and commits high-certainty subsets by unmasking them, with certainty measured by balancing token reliability against external dependency leakage. Empirically, SCCE improves the accuracy–efficiency frontier over local confidence-based unmasking baselines while adding negligible measured per-step overhead in our implementation.