Stability-Aware Self-Training for CLIP under Cross-Modal Anchoring Mismatch
Abstract
Self-training is a practical way to adapt CLIP with limited supervision, but its performance is highly sensitive to pseudo-label reliability. We identify an important failure mode in CLIP adaptation: temporal prediction instability, where a target sample repeatedly flips its prediction across epochs. We provide a perspective based on cross-modal anchoring mismatch: the decision geometry induced by pretrained text anchors can deviate from the local visual cluster structure of the target domain. This mismatch can make ambiguous target samples susceptible to prediction flips during adaptation, whereas target-aligned anchoring produces more stable pseudo-labels. Motivated by this observation, we propose Stability-Aware Self-Training (SAST), centered on Stability-Weighted Image Prototypes (SWIP). SWIP builds class-wise visual anchors from temporally stable target samples to provide more reliable adaptation references. We further add a lightweight text-side refinement to reduce confusion-induced instability, and combine the two streams with an adaptive fusion rule. Across six standard benchmarks and three adaptation settings, SAST consistently outperforms strong baselines. The results suggest that explicitly modeling temporal stability is an effective route to more robust CLIP adaptation under self-training.