SOLA: A Structured Operator Library for Attention in Pretrained Vision Transformers
Enzo Tartaglione ⋅ Vito Paolo Pastore
Abstract
Vision Transformers dominate modern image benchmarks, however, what trained self-attention computes head by head remains unclear: prior work either injects inductive biases at initialization or clusters attention patterns visually without committing to a closed-form computation. We ask whether the softmax$(QK^\top)$ inside a trained head can be replaced, without further training, by a structured closed-form operator while preserving prediction and the residual-stream trajectory. We answer yes across ten ViTs and seven pretrainings (ImageNet-1k, ImageNet-21k FT, MAE, DINO, DINOv2, CLIP, SigLIP), and package the result as \textbf{SOLA}: a Structured Operator Library for Attention with three entries, each gated by a per-head diagnostic that upper bounds substitution error. The entries are a fixed 2D convolution kernel for shift heads, a per-image broadcast row for fixed-target heads, and a per-image rank-$k$ SVD of the attention matrix for near-global heads. A rank-adaptive extension picks $k$ per head from the SVD spectrum and extends structural coverage to every head on every backbone, with task-fidelity controlled by a single spectral threshold $\tau$ that trades compression for downstream fidelity. The diagnostics carry over to cross-attention: every decoder cross-attention head in BLIP captioning has effective rank $\approx 1$ and admits the broadcast substitution. A minimal reproducer that runs the full SOLA pipeline on a single backbone is included in the Supp. Mat. as a demo; the full codebase will be released open-source upon acceptance.
Chat is not available.
Successful Page Load