Margin-based Progress Measures for Grokking
Abstract
Grokking is the phenomenon of delayed generalization in which a neural network fits its training set long before achieving high test accuracy. Identifying progress measures beyond accuracy and loss that track this transition remains an active research problem. We study the evolution of instance margins in the output and input spaces of models that grok. In the output space, we consider the logit margin, the label-free gap between the two largest logits. On correctly classified samples, this margin is tightly linked to cross-entropy, but this relation does not generally hold for misclassified samples, which dominate the held-out set during memorization. Across modular-arithmetic tasks learned by MLPs and LSTMs and a synthetic shape-classification task learned by a ResNet-18, the mean test logit margin remains low during memorization and increases with delayed generalization. The input or embedding-space distance to the decision boundary exhibits similar dynamics, while the rank agreement between the two margins reorganizes through generalization. In the ResNet setting, we further observe that memorization creates a pronounced separation between the train and test margin distributions that progressively disappears as the model groks. These results establish margin dynamics as a simple geometric lens on delayed generalization and connect grokking to distributional generalization at the level of model margins.