Efficient Long-Context Continued Pretraining with Delta Attention
Harsh Singh ⋅ Jeffrey Willette ⋅ Krishna Puvvada
Abstract
Long-context continued pretraining is expensive because dense attention scales quadratically with sequence length. To reduce this cost, sparse attention computes only a subset of query-key interactions, exploiting the fact that attention weights often concentrate on a few tokens. However, introducing sparsity during training alters not only the forward computation but also how gradients propagate in the backward pass. We study the degree to which training can be made sparse while preserving long-context retrieval and book comprehension. We use a variant of Delta Attention, which reuses the difference between dense and sparse attention outputs across nearby queries and lets us control sparsity by changing the spacing between fully computed rows ($\gamma$). We find that a short final phase of dense training restores near-ceiling RULER retrieval across all tested strides and the base model’s LongBook score at moderate sparsity. At 256K context, 500 sparse updates at $\gamma=16$ followed by 50 dense updates yield a 23% reduction in optimizer-update time relative to 550 dense updates. At higher sparsity levels, final performance metrics do not fully recover after the same dense finish. These results indicate attention sparsity comes with a controllable compute–quality knob for long-context training, revealing a regime where training compute can be reduced and quality recovered—and a boundary beyond which that recovery becomes more difficult.
Chat is not available.
Successful Page Load