Vendi Anomaly Scores for Efficient and Accurate Anomaly Detection
Abstract
Anomaly detection (AD) is a critical task in machine learning and scientific discovery. Existing methods that operate on learned embeddings typically define anomalies through local density, isolation heuristics, or distance from a fitted distribution, making them sensitive to neighborhood scale, partitioning choices, or restrictive distributional assumptions. In this work, we introduce a new paradigm by formulating anomaly detection in terms of dataset diversity. We propose the Vendi Anomaly Score (VAS), which detects anomalies by quantifying how much a sample changes the diversity of the dataset, as measured by the Vendi Score, when removed. VAS captures both local redundancy and global data structure without relying on density estimation or partitioning heuristics. Moreover, VAS requires no hyperparameter tuning, instead adapting automatically to the spectral structure of the data. VAS is non-parametric and scales linearly with dataset size. Across 148 benchmark embedding datasets derived from 10 source datasets and 4 embedding architectures, VAS achieves state-of-the-art AD performance. We further validate VAS on large-scale ImageNet experiments, where it remains robust as both contamination rate and dataset scale increase.