On the Information Loss of Multi-Token Prediction: Origin and Solution
Abstract
Multi-token prediction (MTP) accelerates autoregressive decoding through blockwise prediction, but remains vulnerable to error accumulation across sequentially generated blocks. Existing work primarily focuses on improving MTP’s acceleration ratio while largely overlooking the compounding errors that accumulate during block-wise generation. In this work, we study MTP from an information-theoretic perspective and identify the strength of inter-block dependencies as a key factor governing its performance. Building on the cumulative information loss, we theoretically quantify the gap of block-sequence representation. Motivated by the analysis, we propose Dual-MTP, a novel MTP training framework that explicitly models inter-block dependencies from two complementary perspectives, thus reducing the block-sequence representation gap across decoding steps. Specifically, Dual-MTP introduces a dual-perspective training objective with random masking, where each block is jointly trained for forward prediction and masked context reconstruction, encouraging robust block-level representations. Experiments across five model backbones and six tasks spanning math reasoning and GLUE understanding show that Dual-MTP consistently outperforms standard MTP on both quality and efficiency, attaining the highest average accuracy on every backbone with gains of +0.3 to +3.0 points and a 1.4× to 1.9× wall-clock speedup over standard autoregressive decoding.