SpRePE: A Spherical Geometry-Aware Position Embedding scheme for Vision Transformers
Abstract
Vision Transformers are increasingly applied to data defined on the sphere, particularly in physics, meteorology, and related scientific domains. Position embeddings provide the spatial information that attention itself does not encode, making their design central to the expressive power of Transformers. However, most existing position embedding schemes are designed for Cartesian grids and therefore do not naturally handle longitude periodicity, polar singularities, or geodesic relations on spherical domains. This mismatch limits the ability of standard attention to model spherical geometry without specialized architectural modifications. We propose Spherical Reflection Position Embedding (SpRePE), a drop-in position embedding scheme for Vision Transformers on spherical data. SpRePE encodes each absolute spherical position by applying Householder reflections to query and key representations. Although each token is encoded only from its own absolute coordinates, the resulting attention inner products induce an explicit spherical relative-position term with a clear geometric interpretation. This formulation injects sphere-aware relative geometric information into standard attention without constructing pairwise attention-bias matrices, introducing task-specific modules, or modifying the backbone architecture. It also avoids the quadratic overhead of relative position biases and preserves the same asymptotic overhead as RoPE. We evaluate SpRePE on spherical image classification, panoramic depth estimation, and global weather forecasting. Across these tasks, SpRePE achieves competitive performance compared with strong position embedding baselines, with particularly clear gains in settings where spherical geometry is important.