Projection Learning: A Principled Way to Overcome Memorization in Distribution Learning
Abstract
We introduce a framework for learning distributions supported on manifolds based on estimating the projection onto the underlying data manifold. Specifically, we define an objective functional over a class of functions whose minimizer recovers this projection. The proposed framework automatically adapts to the intrinsic dimension and allows for data supported on unions of manifolds with varying dimensions. The resulting objective can be optimized using standard neural network architectures. We present population-level results establishing recovery of the manifold projection and prove finite-sample Wasserstein generalization bounds. Furthermore, we validate the approach empirically across a range of settings. In contrast to widely used generative models such as diffusion models, whose objectives can be minimized by memorizing training samples, our formulation penalizes such behavior. Thus, returning training data does not yield low loss. We also introduce a direct memorization metric to verify that the learned projection generates new samples rather than reproducing the training set.