Training with Delta Attention for Long-Context Adaptation
Harsh Singh ⋅ Jeffrey Willette ⋅ Krishna Puvvada
Abstract
Sparse attention reduces the cost of long-context training, but can compromise a model's ability to retrieve information across long sequences. We extend Delta Attention from inference to training, combining sparse attention at every token with periodic dense anchors. Each anchor's output correction is shared across nearby queries, and backpropagating through it lets non-anchor queries send gradients to keys and values beyond their sparse neighborhoods. The anchor stride controls the compute-quality trade-off without additional parameters or auxiliary losses, while preserving the ability to switch between Delta configurations and dense attention throughout training. On Qwen3.5-9B-Base at 256K context, stride-four Delta achieves a $3.19\times$ attention forward-and-backward speedup over dense-equivalent FlexAttention. A separate stride-64 benchmark achieves a $1.40\times$ training-step speedup. Although sparse training substantially reduces retrieval accuracy at 256K, a short dense continuation closes 94.5% of the gap between continued sparse and fully dense training. After 500 stride-four updates and 50 dense updates, single-needle retrieval accuracy reaches 97.0%, compared with 66.2% for a matched sparse continuation and 98.8% for fully dense training. This schedule uses sparse attention for more than 90% of training updates while recovering near-dense retrieval performance.
Chat is not available.
Successful Page Load