SIGMA: A Sigmoid-Gated Sampler for Test-Time Scaling in Diffusion Language Models
Abstract
Autoregressive language models generate under a fixed left-to-right order, whereas diffusion language models (dLLMs) denoise masked tokens through flexible any-order trajectories that may commit multiple positions per step. This makes dLLMs a distinctive setting for test-time scaling: additional inference compute can diversify both token choices and the order in which positions are committed. Existing samplers do not explicitly allocate these two sources of diversity. Confidence-based decoding is reliable but often redundant, while global temperature sampling and random remasking inject stochasticity without distinguishing useful exploration from unreliable commitments. We propose SIGMA, a training-free sampler that formulates each denoising step as a constrained allocation problem over decoding efficiency, token-level diversity, trajectory-level diversity, and model-internal commitment uncertainty. The resulting objective decomposes joint sampling entropy into token and trajectory terms, yielding a sigmoid-form stochastic gate for position selection and an adaptive per-position temperature for token sampling. Experiments on LLaDA and Dream across math and code benchmarks show that SIGMA improves the accuracy--cost Pareto frontier and can be integrated into existing dLLM test-time scaling pipelines.