Block Parallelism for Efficient Distributed Long-Context Diffusion Language Model Training
Tarun Suresh ⋅ Pranshu Chaturvedi ⋅ Hangoo Kang ⋅ Parth Shroff ⋅ Ishan Khare ⋅ Hermann Kumbong ⋅ Azalia Mirhoseini
Abstract
Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On $16$ H200 GPUs at $256\mathrm{K}$ context, CSBP improves throughput over the best baseline by $\mathbf{1.18\text{--}1.45\times}$ for supervised fine-tuning and $\mathbf{1.27\text{--}1.33\times}$ for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches $\mathbf{1.61\times}$ at $512\mathrm{K}$. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by $\mathbf{2.48\times}$ at $512\mathrm{K}$ and $\mathbf{7.59\times}$ at $1\mathrm{M}$. In matched $12$-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite.
Chat is not available.
Successful Page Load