Prompt Ensemble Image Purification for Test-time Adversarial Robustness of CLIP
Abstract
Vision-Language Models like CLIP exhibit remarkable zero-shot capabilities but remain vulnerable to adversarial attacks. Existing defenses, such as adversarial fine-tuning or test-time defense, either incur high computational costs that lead to cause catastrophic forgetting, or struggle against strong adversarial attacks. In this work, we reveal that adversarial perturbations are highly overfitted to the decision boundary of a single canonical text prompt. By introducing several semantic prompt variants, we identify a \textit{Prompt Sensitivity Gap}: adversarial examples exhibit significantly higher prediction variance across prompts compared to benign images. Motivated by this insight, we propose Prompt Ensemble Image Purification (PEIP), an efficient test-time defense framework. PEIP features a dual-objective purification loop that jointly suppresses prediction variance to dismantle adversarial alignment and reinforces the most responsive class across prompt variants to facilitate semantic recovery. Furthermore, to accelerate inference, we introduce a statistically calibrated Early Exit mechanism that bypasses benign images based on their initial prompt sensitivity. Extensive experiments across 16 classification benchmarks, multiple CLIP architectures, and CLIP-based zero-shot semantic segmentation tasks demonstrate that PEIP achieves state-of-the-art adversarial robustness while preserving exact zero-shot accuracy and high efficiency.