Native Geometry vs. Linear Recoverability: How Alignment Changes Platonic Scaling
Abstract
Do larger models learn increasingly similar representations of the same world? The Platonic Representation Hypothesis (PRH) predicts that representations should become more similar as models scale, but the answer depends on what we mean by "similar.'' Here, we compare two kinds of similarity: the native cosine geometry of the representations, and the correspondence that can be revealed by learning a supervised linear map between them. Using matched embeddings of the same astronomical objects from the Legacy Survey and Hyper Suprime-Cam (HSC) across five model families, we find that linear alignment strengthens the relationship between model size and cross-dataset neighbourhood correspondence compared with native cosine geometry. This effect appears in all five model families and remains when both representations are reduced to the same fixed dimension of (256), ruling out wider embeddings simply giving the alignment more dimensions to work with. However, the learned linear maps are far from simple similarity transformations and distort different directions by different amounts, including on held-out local data directions. Sparse Autoencoder (SAE) and Block-Sparse Featuriser (BSF) representations also provide no advantage over the matched dense-alignment baseline. Overall, while model scale systematically improves linear recoverability across datasets, it does not support that native representations are converging toward an isometric or shared metric geometry.