LayerScope: A Layerwise Audit of Video and Multimodal Learned Representations
Abstract
Video and multimodal models are typically evaluated through final-layer representations or model-default outputs, although intermediate layers and combinations of layers are also used. Determining which representations are most useful for downstream performance can require labeled data, repeated downstream experiments, and substantial computation across models and layers. To address these limitations, we introduce LayerScope, a layerwise framework that audits learned representations without task labels. LayerScope characterizes representations using complementary local, global, distributional, and correspondence-sensitive geometric metrics, enabling systematic comparison of layerwise organization within and across models. We evaluate seven architecturally diverse video and multimodal models on classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs, but no single geometric metric consistently predicts downstream performance across tasks, as different tasks exhibit distinct geometric signatures. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that correspondence-sensitive metrics explain retrieval better than distributional distance alone. These results show that layerwise geometry provides a useful, label-free diagnostic of representation organization that complements downstream evaluation. LayerScope offers a framework for comparing representations across models and layers, enabling a more systematic analysis of learned representations in video and multimodal settings.