PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models
Xiaocheng Lu ⋅ Shuhan Guo ⋅ Ziyue Ma ⋅ Jie ZHANG ⋅ Jian Liu ⋅ Jingcai Guo ⋅ Haoxuan Che ⋅ Song Guo
Abstract
Diffusion large language models (dLLMs) such as LLaDA and Dream now rival autoregressive LLMs in quality while retaining native parallel decoding. To deploy them efficiently, block-wise decoding partitions generation into blocks of size $B$ and commits a fraction of each block before moving on, but $B$ conflates two roles: \emph{how far the model can look ahead} and \emph{how many tokens get committed per step}. Recent accelerators relax this trade-off with indirect heuristics, yet the underlying difficulty is already exposed at every forward pass by the model's per-step confidence. Specifically, the in-window confidence profile exhibits a context-dependent \emph{cliff} (a high plateau, a sigmoidal transition, and a residual floor) that directly encodes how far ahead is safe to look. We propose \textbf{PACE-dLLM}, a decoder that handles the two roles independently. At each step, PACE-dLLM fits the cliff in closed form on the current window's confidences and reads off the next prediction horizon at the cliff's saturation point; a separate confidence threshold decides which tokens are committed. Both rules read the same signal, add a single hyperparameter, and admit a strict-dominance guarantee over fixed-block decoding. Extensive experiments across four reasoning and code benchmarks on two open-source dLLM backbones demonstrate that PACE-dLLM achieves the best average accuracy at a roughly $\mathbf{4.5\times}$ wall-clock speedup over the unaccelerated semi-AR baseline, pushing the Pareto frontier outward.
Chat is not available.
Successful Page Load