Representation geometry recovers relatedness, not phylogeny
Saket Atreya ⋅ Paras Chopra
Abstract
Frozen self-supervised encoders place related sources close together, and this is often read as evidence that they have captured the process that produced them. We show that proximity and branching order come apart, and that they must. Across five modalities (text, speech, protein, vision, single-cell), measured against externally known gold structure, distances are faithful: quartet accuracy runs $0.55$--$0.86$ against a chance level of $1/3$, and nearest-neighbour purity exceeds its per-domain chance everywhere. Topology is not recoverable: exact neighbour-joining recovery is $0.00$--$0.22$, while the same pipeline recovers $0.97$ of the deep clades of a synthetic additive tree. The failure is therefore a property of the data, not of the estimator. The dissociation is predictable rather than incidental. Neighbour-joining is exact only within a tolerance set by the shortest internal branch of the true tree, whereas the correlation between distance matrices degrades only quadratically in the same perturbation. Any process that makes a population non-additive, such as contact between lineages or horizontal transfer, therefore destroys the tree at perturbation scales that leave distances close to intact. A synthetic dial confirms the ordering, and curvature, the training objective, cloud higher moments and monotone distance corrections are each excluded as alternative explanations. What can be read at all depends on which group is quotiented out: the aligned centroid norm tracks a corpus-level nuisance, so a Euclidean reading of 36 languages gives a null partial correlation where cosine gives $+0.43$, and a collapsed final layer reports no signal where mid-stack reports a clear one. We give label-free diagnostics for such choices. Proximity in a learned representation is evidence of relatedness, not of descent.
Chat is not available.
Successful Page Load