TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
Abstract
Adapting CLIP for anomaly detection on unseen objects has shown strong potential in a zero-shot manner. Existing methods typically rely on a single textual space to align with visual semantics across diverse objects and domains. The indiscriminate alignment hinders the model from accurately capturing anomaly semantics. We propose TokenCLIP, a token-wise dynamic framework that aligns each visual token with learnt textual subspaces corresponding to its visual characteristics. However, explicitly assigning a unique learnable textual space to each token is computationally intractable and prone to insufficient optimization. We instead expand the token-agnostic textual space into a set of orthogonal subspaces, and then dynamically assign each token to a subspace combination guided by semantic affinity, which jointly supports customized and efficient token-wise adaptation. To this end, we formulate dynamic alignment as an optimal transport problem, where all visual tokens in an image are transported to textual subspaces under the cross-modal cost matrix. \textbf{The marginal constraint and minimal cost objective of OT ensure sufficient optimization across subspaces and encourage them to focus on different semantics.} Solving the problem yields a transport plan that adaptively assigns each token to semantically relevant subspaces. Extensive experiments show that the textual subspaces naturally specialize in different semantics, such as foreground and background, to promote fine-grained anomaly learning. The comparison between baselines shows the superiority of TokenCLIP.