Auditing Geometric Change in Grokking: Centroid Configuration versus Within-Class Geometry
Felix Berg ⋅ Ulrik Unneberg
Abstract
Geometric statistics often change when networks grok, but a global measurement does not show whether class arrangement or geometry within classes changed. We separate these effects by writing each activation as a class centroid plus a within-class residual, then intervening on and recombining the two components. In nine $\mathbb{Z}_{23}$-addition MLPs, changing only the class-centroid configuration moves the finite-sample MST-based PH-dimension estimate by more than an order of magnitude across the intervention sweep. TwoNN, in contrast, becomes invariant once the neighbours it uses lie within classes. Checkpoint recombination gives the same qualitative contrast over training: in every seed, the centroid contribution to PH-dimension exceeds the residual contribution under the symmetric endpoint allocation, whereas TwoNN is residual-dominated. Exact numerical shares depend on how the factor interaction is divided. The difference follows from estimator support: TwoNN is local, while the MST keeps cross-class connectors. A modular-division Transformer shows that the attribution also depends on the chosen training interval.
Chat is not available.
Successful Page Load