A Geometric Invariant of Attention Identifies Head Specialization from Weights Alone
Miguel Pedra Bento
Abstract
Transformers are built with well-structured continuous symmetries in the attention mechanism. We derive $\rho$, a scalar invariant that measures each head's structural capacity to couple to rotary positional encodings. This framework allows us to analyze heads without forward passes, opening a new research direction into data-independent analysis of modern LLMs. Empirically, we find that $\rho$ can be used to identify induction heads, the most-studied circuit type in mechanistic interpretability, with Spearman correlation $r = -0.70$. Furthermore, we use $\rho$ to apply YaRN to only the top-50\% positional heads, achieving better context performance than uniform application, thus confirming that the invariant correctly differentiates content-dominated from position-dominated heads.
Chat is not available.
Successful Page Load