One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
Abstract
Universal visual anomaly detection (AD) aims to identify anomalous images and segment anomalous regions towards open and dynamic scenarios, typically adhering to zero- and few-shot paradigms without dataset-specific fine-tuning. Recently, the field has seen significant progress driven by the integration of vision-language foundation models. However, we observe that current methods often struggle with laborious prompt engineering, elaborate adaptation modules, and complex training strategies that ultimately constrain their flexibility and generalizability. In this paper, we rethink the fundamental mechanism of vision-language models for AD and present UniADet, an embarrassingly simple, effective and general framework for Universal vision Anomaly Detection. Our approach is built on two key insights: First, we reveal that the primary function of the language encoder is merely to derive decision weights and we demonstrate that it is unnecessary, as these weights can be learned more directly and efficiently. Second, to resolve learning conflicts arising from disparate feature manifolds, we introduce a systematic decoupling strategy that learns independent weights across both distinct tasks (classification vs. segmentation) and hierarchical features. UniADet is highly simple and efficient, learning only decoupled weights and requiring 0.02M learnable parameters. It is inherently versatile, adapting seamlessly to various foundation models (\eg, CLIP, DINOv2, and DINOv3). Extensive evaluations of 14 real-world benchmarks in the industrial and medical domains demonstrate that UniADet not only exceeds state-of-the-art zero/few-shot methods by a substantial margin, but also outperforms full-shot AD methods for the first time. This empirical evidence reveals the profound potential of language-free frameworks to redefine the boundaries of visual anomaly detection. The code and models will be made publicly available.