Source and Canvas Readouts Redistribute with Lexical Commitment in Masked Diffusion Language Models
Zhouyang Jin ⋅ Bo Xue ⋅ Jiaxin Ding ⋅ Luoyi Fu ⋅ Xinbing Wang
Abstract
Masked diffusion language models repeatedly revise a fixed canvas whose composition changes as unresolved MASK positions acquire discrete tokenizer identities. We index this process by \emph{lexical commitment}, $k$, the number of canvas positions fixed to tokenizer symbols under a deterministic policy that commits one token per step. Using paired source--canvas measurements from the same forward passes, we identify a replicable, commitment-indexed redistribution between two prespecified pooled sentence-geometry readouts in Dream and LLaDA. This regional endpoint contrast transfers across STS14 and SemRel more consistently than a secondary, checkpoint-specific coarse-timing contrast. Targeted interventions provide an independent functional anchor: late hidden states transfer token support across texts matched by position and token, while rolling back the prediction-aligned state immediately before the final commitment produces excess token-change rates of 15.9 and 3.05 percentage points in Dream and LLaDA, respectively. On STS13, an early source-derived direction affects later lexical and terminal outcomes in Dream; first-state interventions propagate beyond their two possible entry sites in both checkpoints; and natural-state restoration recovers deterministic replay in each. Together, these results motivate commitment-aware reporting: a fixed physical support set does not by itself specify an intermediate DLM readout, which additionally requires its commitment state and pooling region.
Chat is not available.
Successful Page Load