SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Zhiwei Li ⋅ Lei Zhu ⋅ Hao Gu ⋅ Xiang Hu ⋅ Yan Wang ⋅ Haitao Mi ⋅ Sirui Han ⋅ Leo Liang ⋅ Zhijiang Guo
Abstract
Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually rely on a lightweight selector to score context units followed by a hard Top-$K$ selection, which blocks gradients from the language modeling loss. As a result, these methods commonly resort to distilling layer-wise dense attention distributions. While this approach encourages the selector to rank context units according to dense attention weights in the original model, such a ranking is not directly aligned with their impact on the model's final predictions under a fixed attention budget (ie, the number of attended context units per query), which can waste the limited attention budget on less useful units. To address this ranking misalignment, we propose **S**imple **A**ttention **S**parsification (SAS), a gated sparse attention mechanism that optimizes context ranking end-to-end with the language modeling loss. The key idea is to inject the selector's continuous scores into the attention logits during training, allowing the language modeling loss to update the selector through standard backpropagation. We identify several choices that are crucial to make this simple design work well in practice: placing the gate inside the attention $\operatorname{softmax}$ in log form, keeping all context units active during training so each of them can receive gradients, and preserving continuous selector scores so the model learns relative priorities rather than only hard selections. Since a naive implementation that explicitly materializes the full attention matrix would cause prohibitive memory overhead for long sequences, we implement a memory-efficient Triton kernel that integrates the design into a FlashAttention-style computation. SAS consistently outperforms trainable sparse-attention across attention budgets, with especially large accuracy gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
Chat is not available.
Successful Page Load