The Fault in Our Metrics: Revisiting Generative Model Evaluation with Chamfer Distance
Abstract
Reliable evaluation is essential for progress in image generation, yet widely used metrics rely on assumptions that are misaligned with the geometry of real data. We revisit two dominant families of metrics: Fréchet-based metrics such as FID, and precision/recall-style metrics based on KNN balls. We show that FID reduces complex feature distributions to a Gaussian moment summary, making it blind to differences beyond mean and covariance. We further show that KNN-ball metrics construct inflated isotropic support estimates in high-dimensional feature spaces, causing binary membership queries to accept off-manifold samples and making recall unstable under standard finite-sample evaluations. We propose Chamfer distance as a simple alternative: instead of constructing inflated support estimates, it directly measures cross-set nearest-neighbor distances between real and generated samples. The resulting one-sided distances provide interpretable signals for fidelity and coverage, while reducing sensitivity to the neighborhood hyperparameter and remaining computationally efficient.