A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of $\mathrm{O}(2)$
Tikun Ong ⋅ Georg Bökman
Abstract
We introduce a family of vision transformers equivariant to arbitrary discrete subgroups of $\mathrm{O}(2)$, providing a unified framework that generalizes prior flipping- and $D_4$-equivariant architectures, together with a novel family of ViTs operating on hexagonal patches, compatible with six-fold rotational symmetries. The construction of equivariant transformer components is accompanied by expressivity guarantees: we show that whenever $H \le G$, the class of $G$-equivariant ViTs embeds naturally into the class of $H$-equivariant ViTs. We also prove that our equivariant self-attention layer realizes every $G$-equivariant map representable by ordinary self-attention in the single-head case. The resulting models are evaluated on the PatternNet aerial image dataset across subgroups of $D_4$ and $D_6$. While equivariance helps in artificially data-scarce regimes, the accuracy gap is mostly closed by more training examples or data augmentation. It is observed that the class token tends to overweight the trivial irrep component, which may explain why larger symmetry groups do not necessarily help on this classification task.
Chat is not available.
Successful Page Load