Visualizing Dense Neural Representations
Abstract
To study the features learned by recent dense-feature models qualitatively, representations are usually mapped onto the first three principal components and then to RGB values. However, even between models that are in the same ballpark for downstream task performance, the noisiness of these visualizations varies substantially, seemingly indicating large differences in feature quality. Here, we propose a simple, completely unsupervised tweak to the PCA objective, which identifies components with much less high-frequency noise, showing that such components can be extracted from all models to some extent. Qualitatively, these visualizations explain why differences in task performance between models are relatively modest. We quantitatively show that the resulting components align substantially better with dimensions that are relevant for semantic downstream tasks, demonstrating that they paint a more complete picture of a model's feature space.