No Equivariant Architecture Covers All Equivariant Attention
Tikun Ong
Abstract
We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$ can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: equivariant maps representable by unconstrained MHSA form a union of extremely many Zariski-irreducible components in the reduced parameter space, and any single architecture covers at most one. For $G=D_4$ acting on $C$ (even) copies of the regular representation as the token feature space, we show that there are at least $\Omega(C^{64})$ components for eight attention heads.
Chat is not available.
Successful Page Load