OrbitLoRA: Learning Rotation-Aware Low-Data Adaptation of Vision Foundation Models
Abstract
Pretrained vision transformers provide strong semantic representations, but their token embeddings are not designed to transform predictably under image rotations. This limits their direct adaptation to domains such as satellite imagery, histopathology, microscopy, and other scientific imaging settings, where in-plane orientation is often arbitrary and no canonical pose exists. Existing equivariant adaptation methods for pretrained models typically rely on frame averaging or canonicalization, which can be effective on synthetically rotated natural images but do not teach the adapted model an internal rotation-aware representation. We introduce OrbitLoRA, a lightweight orbit-aware adapter for frozen visual foundation models. OrbitLoRA combines standard LoRA with a typed residual branch that maintains scalar and vector fields over ViT patch tokens. The typed branch is initialized from local steerable image features and harmonic coordinate fields, updated at selected transformer depths, and injected into the frozen token stream through zero-initialized residual gates. During finetuning, paired rotated views supervise the typed state through a transport-and-rotation consistency loss, while the semantic output is trained to remain invariant for the downstream task. This design preserves the rich features of pretrained ViTs while adding a learnable rotation-symmetry prior. Across classification and segmentation tasks, and across DINOv2-, CLIP-, SAM-, and LLaVA-style backbones, OrbitLoRA matches strong LoRA-with-augmentation baselines in high-data regimes and substantially improves adaptation in low-data regimes, where learning the symmetry prior is most valuable.