Learning Symmetry from Limited Data Augmentation with Transformers
Abstract
Training transformer-based architectures with data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we investigate the ability of a vanilla transformer to learn a wide range of symmetries through limited data augmentation. We identify the following hierarchy: performance is weakest for non-angle-preserving symmetries and within the angle-preserving family, base subgroups such as translation, rotation, and scale are easier than compositional groups that combine them. Motivated by this observation, we perform a structural analysis of the trained models and identify interpretable mechanisms that induce invariance to each base angle-preserving symmetry.