Training Selects Between Loss-Degenerate Representational Geometries in a Toy Model of Superposition
Cliff Schmidt
Abstract
Neural networks can realize essentially the same task loss with qualitatively different internal geometries; we ask which geometry a specified training procedure selects. In an overcomplete toy model of superposition ($n$ features, two hidden units, ReLU reconstruction), steeply decaying feature importance admits a monosemantic geometry that drops low-importance features and an S-side superposed geometry at the numerical loss floor ($\sim 10^{-9}$). Under Adam training from random initialization with a sparsity curriculum, none of 8 seeds reaches the monosemantic side in any of six steep-importance cells (sizes $n \in \{144,200,256\}$ $\times$ importance decay $\rho \in \{0.7,0.8\}$; 48 runs total). The constructed monosemantic solution is nevertheless stable under the final target-task dynamics: warm-started networks remain on the monosemantic side (alignment $\bar D_{\mathrm{end}}=0.974$–$0.997$, all sizes). The stable-but-unreached pattern disappears at $\rho=0.9$ (7–8/8 seeds reach the monosemantic side, all sizes), and a non-degenerate uniform-importance control geometry is recovered cleanly, arguing against a generic convergence failure. Prior work established that loss-matched monosemantic and polysemantic solutions can coexist (Jermyn et al., 2022); our contribution is an importance-axis map of this selection, a constructed-solution stability test, and finite-size consistency checks. The result suggests that the geometry exhibited by a trained network can depend on training dynamics even when task loss does not distinguish the alternatives.
Chat is not available.
Successful Page Load