Belief-State Geometry beyond Next-Token Prediction
Abstract
What computational structure are we building into large language models when we train them on masked-diffusion objectives? Prior work showed that a transformer trained on next-token prediction represented belief-state geometry linearly in its residual stream (Shai et al., 2024). However, next-token prediction can be considered a special case of masked diffusion, limiting masks to after the current token. We extend this analysis to the full family of masked-diffusion objectives; initially deriving the closed-form family from the mess3 process, then measuring in two ways: on small transformers, as in the original work, and on pretrained 7-8B checkpoints. We find that the measurement matches the closed-form solution closely (held-out R^2 is 0.962-0.995), with the masked-diffusion read-out surpassing the upper bound score achievable from past tokens alone. We also investigate belief state geometry at each layer in the network, showing that about half of the fading in later layers is a log-coordinate change, recoverable by refitting the probe. Finally, we show the pretrained checkpoints alone contain this geometry innately, even before finetuning. Overall, we find a single law defining contraction, with rate fixed by the process itself - which connects next-token prediction to the wider masked-diffusion family.