Rethinking Memorization–Generalization Trade-Off in Generative Models
Abstract
Standard generative models face a memorization--generalization trade-off, that is avoiding memorization is considered necessary for generalization. In supervised learning, however, recent studies show that highly overparametrized models can memorize (interpolate) training data while still generalizing well, a phenomenon known as benign overfitting. Motivated by this, we investigate whether generative models can similarly bypass the trade-off and generalize while interpolating. Generative models should minimize the distance between the true data distribution and the distribution induced by mapping the full latent distribution through the generator. But since the true distribution is inaccessible, existing models instead minimize empirical risk with respect to the training distribution. Under this standard formulation, exact empirical-risk minimization forces the generator to produce only training samples. To address this issue, we consider an alternative empirical risk based on presampled latent variables. Our experiments demonstrate that the resulting presampled-latent version of the flow matching model exhibits benign overfitting on standard image benchmarks, such as MNIST and CIFAR-10. As a theoretical proof of concept, we recast generative modeling as a regression problem and extend existing benign overfitting theory to our setting. Together, these results establish, for the first time, that benign overfitting can occur in generative models.