Spectral Degeneration of Softmax Attention under Isotropic Score Geometry
Mengda Li ⋅ Jianfeng Yao
Abstract
We study the spectral structure of the attention matrix in Transformer models through an idealized spectral model for the pre-softmax score matrix. Motivated by the rotational isotropy of query and key representations, we replace the random token-space orientation of the score matrix by its canonical singular-value representative $\Sigma$, and assume that $\Sigma$ has a logarithmic spike-bulk separation. We provide empirical diagnostics consistent with this model on pretrained vision and language transformers: token-space singular vectors exhibit near-Haar behavior, and score spectra are dominated by a small number of effective directions. We then prove that the row-wise softmax map transforms this structure into a degenerate attention matrix: dominant modes are preserved, whereas the bulk spectrum converges weakly to $\delta_0$. This provides a unified theoretical explanation for why attention layers can act as low-rank information extractors: they preserve a small number of dominant token directions while compressing the remaining bulk modes.
Chat is not available.
Successful Page Load