How Wide Is Attention? Measuring Its Effective Width in Tokens
Karen Mosoyan ⋅ Henry Ndubuaku ⋅ Jakub Mroz ⋅ Noah Cylich ⋅ Roman Shemet ⋅ Parkirat Sandhu ⋅ Satyajit Kumar ⋅ Justin H Lee
Abstract
Global softmax attention provides language models with a runtime memory whose nominal width is the context length T. We study how wide attention is when interpreted as an MLP whose up projection's rows are the keys and down projection's rows are the values. We define $M_\varepsilon$, which is the minimum number of real KV cache tokens which can reconstruct an attention head's outputs over a query distribution within relative error $\varepsilon$. We estimate $M_\varepsilon$ with a greedy algorithm and evaluate it across dense Transformers, hybrid architectures, and attention-only SAN models at context lengths up to 128k. We find that effective width generally grows with context length, but strongly decelerates at longer lengths. Comparing natural and entropy-matched isotropic queries suggests roles ranging from MLP-like pointwise computation in SAN to broader query coverage in hybrid models. In Qwen3-8B, budgets calibrated from effective width closely preserve language-modeling loss but lose full attention's perfect recall on our needle-in-a-haystack diagnostic.
Chat is not available.
Successful Page Load