CURE: Counterfactual Unsafe-token Re-masking for Diffusion Large Language Model Test-time Alignment
Abstract
Diffusion language models (DLMs) generate text through iterative denoising, enabling parallel decoding, bidirectional conditioning, and editable intermediate states. Under harmful prompts, however, intermediate denoising states can contain risk-inducing tokens that are seemingly benign but increase the probability of an unsafe final response by shaping how the remaining masked positions are completed. Since committed tokens in DLMs can still be re-masked and rewritten, safety control can be applied during denoising by removing a small set of risk-inducing tokens before they drive subsequent generation toward unsafe responses. Existing DLM defenses mainly rely on model-level alignment, trajectory-level detection, or block-level repair, often requiring training cost, extra inference passes, or coarse regeneration. We propose CURE, a counterfactual value-driven test-time alignment method that keeps the base DLM fixed and selectively re-masks risk-inducing tokens. CURE trains a time-conditioned safety value model to estimate whether a partially denoised state will lead to an unsafe final response. During inference, CURE constructs counterfactual readout views that remove or isolate targeted tokens, estimates their contribution to future unsafe response, and re-masks only high-risk tokens for later rewriting. Across three DLMs and five jailbreak benchmarks, CURE reduces macro-average ASR from 47.20\% to 1.36\%, preserves utility on MMLU and GSM8K, and adds only 0.84 extra TFLOPs, achieving a strong safety-utility-efficiency trade-off. Code is available at \url{https://anonymous.4open.science/r/CURE}.