Winning the Symmetry Lottery: Learning Invariance from Data Augmentations with Transformers
Abstract
Training transformer-based architectures with data augmentation has become an increasingly popular approach for equivariant machine learning. Despite its empirical success, the interplay between the transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we investigate the ability of a vanilla transformer to learn a wide range of symmetries through limited data augmentation. We identify that it performs strongly for common angle-preserving symmetries, while it struggles with non-angle-preserving ones. Within the angle-preserving family, we further find that base subgroups such as translation, rotation, and scale are learned more easily and accurately than compositional groups that combine them. Motivated by this observation, we perform a structural analysis of the trained models and identify interpretable mechanisms that induce invariance to each base angle-preserving symmetry. These findings support what we coin winning the Symmetry Lottery: the transformer architecture is aligned with learning mechanisms invariant to symmetries that happen to be prevalent in scientific domains.