Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces
Andrew Bond ⋅ Ege E Özlü ⋅ Tuna Çimen ⋅ Ilkin U Melanlioglu ⋅ Tolga Birdal ⋅ Erkut Erdem ⋅ Aykut Erdem
Abstract
Methods that operate on Vision Transformer features almost universally rely on Euclidean distance to judge feature similarity. Yet Euclidean distance weights every direction in feature space equally, ignoring that downstream tasks concentrate their sensitivity on a low-dimensional submanifold. The metric that respects this structure is the pullback Riemannian metric $g(F) = J(F)^\top J(F)$ induced by the task decoder's Jacobian, but materializing it is prohibitively expensive; a single ViT-B/14 probe point requires storing $4\times 10^{10}$ scalars, and per-token scoring provably fails when task-sensitive directions are spread across tokens. We show that whether a low-rank approximation of $g(F)$ is \emph{learnable} at all depends on a precise structural property of the model-decoder pair, which we characterize via a single scalar diagnostic $\kappa_{cap}(r)$, the captured-energy fraction of the top-$r$ singular directions of $J$, computable matrix-free in $\mathcal{O}(m+rq)$ Jacobian-vector products. When the diagnostic licenses it, we train the \emph{Spectral Pullback Network} (SPN) on directions found by randomized power iteration on $J^\top J$, and distill the resulting per-token signal into a 310K-parameter feature-only \emph{conformal head}, never materializing $J$. When the original spectrum is too rich for any practical rank-$r$ budget, a VAE reparameterization of the decoder recovers a latent space in which the same machinery applies, converting intractable dense pipelines into tractable ones. Across DPT, DINOv2, CLIP, and VGGT backbones the diagnostic correctly predicts which architecture succeeds: the conformal head reaches Spearman $\rho=0.998$ against the Jacobian-derived target on DINOv2 CLS, and geometric token pruning yields a 25\% relative improvement over ToMe at aggressive prune ratios on DPT depth. These results suggest that learning \emph{which} directions matter is often more valuable than learning better features.
Chat is not available.
Successful Page Load