Structure–Function Relationships in Vision Models
Raghav Jain ⋅ Meenakshi Khosla
Abstract
Can we predict what a system does from the structure of what it represents? We introduce a meta-modeling framework that uses internal representations to predict behavior across more than 1,200 pretrained vision models. To account for arbitrary activation coordinates, we construct symmetry-aware descriptions using representational dissimilarity matrices, Centered Kernel Alignment profiles, and Procrustes-barycenter coordinates. We learn structure–function relationships on CNNs and test whether they generalize to unseen architecture families. We also examine how stimulus selection affects prediction and whether representations of in-distribution images predict behavior under distribution shift. A few hundred informative images predict ImageNet performance in held-out non-CNNs with test R$^2$ approaching 0.9. The same representations also predict out-of-distribution robustness. These results establish population-level meta-modeling as an approach for identifying structure–function relationships that generalize across architecture families and behavioral regimes.
Chat is not available.
Successful Page Load