Step-dLLM: Adaptive Step-aware Sparse Attention for Efficient Diffusion LLM Inference
Zhichen Zeng ⋅ Xichong zhang ⋅ Junpan Wu ⋅ Yifei Zuo ⋅ Chi-Chih Chang ⋅ Jiayi Wang ⋅ Maohua Nie ⋅ Ji Liu ⋅ Ang Li ⋅ Banghua Zhu
Abstract
Diffusion large language models (dLLMs) generate tokens in parallel via iterative denoising, but their bidirectional attention prevents the KV caching used by autoregressive models and recomputes the full attention map at every step, incurring quadratic per-step cost. Existing sparse attention methods for dLLMs compute a sparse mask once at the first denoising step and reuse it throughout, while allocating the sparsity budget uniformly across heads. We argue both choices ignore key structural properties of attention in dLLMs: empirically, attention patterns are only locally consistent across steps and shift abruptly at certain transitions, and different heads concentrate their attention mass to very different extents. To this end, we design Step-dLLM, which periodically refreshes a block-sparse mask at anchor steps along the denoising trajectory and reuses it within each locally consistent interval, while a head-adaptive global budget assigns more blocks to heads with more concentrated attention. Backed by a fused IO-aware Triton kernel for block-sparse attention, Step-dLLM applies to pre-trained LLaDA-1.5 and Dream-7B-Instruct without retraining, achieving a $4.2\times$ kernel-level speedup over dense FlashAttention at 16K context with 90% sparsity (up to $8.5\times$ at 95%), and achieves better model accuracy and lower latency than prior methods.
Chat is not available.
Successful Page Load