More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
Noam Elata ⋅ Itay Lamprecht ⋅ Mikey Shechter ⋅ Daniel Ohayon ⋅ Itay Hubara ⋅ Daniel Soudry
Abstract
Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while value heads preserve capacity at no additional cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple GPU-agnostic sparse attention method designed to isolate sparsity's effect on decoding-optimized architectures. We formalize the benefits of this asymmetry theoretically and validate it empirically, achieving end-to-end decoding speedups exceeding $2\times$ at long contexts across multiple model scales and benchmarks. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants while delivering substantial efficiency gains. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture at a small to negligible cost to quality, enabling practitioners to benefit from our approach without costly retraining.
Chat is not available.
Successful Page Load