Memorization in image classification: test set behavior and scaling laws
Abstract
Deep neural networks are capable of both learning generalizing solutions on natural tasks and memorizing arbitrary mappings. We argue that understanding the former requires understanding the latter, which is much less studied. In what ways do networks that generalize or memorize differ in their behavior or their training dynamics? Do memorizing networks simply function as a lookup table where input-output mappings are stored independently? We investigate these questions in the context of CNNs trained on image classification tasks with randomized labels or pixels. We first study the behavior of such memorizing networks on unseen images. We find that independently initialized networks trained on the same random labels eventually agree on the memorized training examples but not on the test examples: they learn different predictive functions, in contrast with generalizing networks. We then examine the effect of dataset size on memorization through the required network capacity and training time to interpolate the training set. Counter-intuitively, we find that the number of presentations of a given example required for complete memorization can decrease with dataset size even when input-output pairs share no common structure. Although memorizing networks do not learn a universal solution, they exhibit interesting behaviors which call for further investigations into the underlying mechanisms.