Benefit of Structured Representations: When Equivariance Helps in Image Classification
Florian Hölzl ⋅ Daniel Rueckert ⋅ Georgios Kaissis
Abstract
Equivariant neural networks promise improved representation learning through structural guarantees. However, mainstream vision encoders rarely incorporate equivariance for most geometric symmetries. Images are a particularly interesting modality with many such symmetries and extensive datasets from which to learn this structure. In this work, we investigate when vision encoders benefit from geometric inductive biases. We study the discrete rotation group of quarter-turns ($C_4$) which acts losslessly without interpolation or boundary artifacts. Combined with image classification as a controlled invariant objective, this allows for a principled empirical investigation into the role of structural guarantees. We introduce ConvNeXt-C4 as an rotation equivariant ConvNeXt with dimension-equivalent yet orientation-structured representations. Training from scratch, our experiments show, that the C4 variant outperforms its base model across all tasks while operating with $6.4\times$ fewer parameters and $1.6\times$ fewer FLOPs. Its performance advantage increases log-linearly with decreasing training set size, with pre-training providing a similar improvement across both model variants. An analysis of the classifier head reveals that models actively use orientation information from the representations' structure, improving performance by $4.9\%$ across all datasets. Our results underscore the benefits of equivariance beyond sample-efficiency and that structure remains an important component not only under large-scale pre-training but possibly also for more complex tasks outside of image classification.
Chat is not available.
Successful Page Load