Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
Sheir A. Zaheer ⋅ Jihwan Moon ⋅ Chan Youn Park
Abstract
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (HW-REViTs) are available at (\emph{removed for anonymity}).
Chat is not available.
Successful Page Load