Latent Algorithmic Structure Precedes Grokking: A Mechanistic Study of ReLU MLPs on Modular Arithmetic
Anand Swaroop
Abstract
Grokking on modular addition has been characterized by sinusoidal representations in transformers and MLPs. We find that ReLU MLPs in our setting instead learn near-binary square-wave input weights, with intermediate values appearing only near sign-change boundaries, and output weights whose dominant Fourier phases satisfy $\phi_{\text{out}} \approx \phi_a + \phi_b$. These relations persist under label noise and even in models that do not grok. Using DFT-extracted frequencies and phases to replace learned weights with ideal square/cosine waves yields **95.8\% accuracy** from noisy models that achieve only 0.40\%. This suggests that grokking sharpens an algorithm already encoded during memorization.
Chat is not available.
Successful Page Load