Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective On Robustness
Abstract
Positional embeddings (PEs) in Vision Transformers (ViTs) are known to affect performance and robustness, but their role in shaping internal spatial representations is not well understood. We study how different PEs influence ViT representational geometry and how these changes relate to robustness under content-disrupting distribution shifts. We introduce the Spatial Similarity Distance Correlation (SSDC) to quantify spatial structure in token representations. ViTs trained without PEs still develop non-trivial spatial structure, but it is content-driven and collapses under token permutation. All PEs considered (learned absolute, sinusoidal, and rotary) instead induce a consistent shift toward index-anchored spatial organization, and the resulting representations remain stable under content-disrupting perturbations, yielding substantially improved robustness. Different PEs produce distinct depth-wise trajectories, yet their robustness is largely similar, with only secondary variation across schemes -- suggesting robustness depends more on having a stable positional reference frame than on the specific encoding mechanism. These results give a geometric account of how positional encodings shape internal representations, with implications for future encoding designs.