[Re] Boosting the Visual Interpretability of CLIP via Adversarial Fine-Tuning
Anton Nuzhdin ⋅ Andrei Gruzitski ⋅ Balázs Egyed ⋅ Aleksandr Raudvee
Abstract
This paper presents a reproducibility study of "Boosting the Visual Interpretability of CLIP via Adversarial Fine-Tuning" by Gong et al. (2025), published at ICLR 2025, which proposes an unsupervised adversarial fine-tuning (AFT) method with norm regularisation to enhance the visual interpretability of CLIP's image encoder. We attempt to reproduce the key claims regarding improved saliency map quality, increased concept alignment, transferability to out-of-distribution datasets, and the trade-off with zero-shot accuracy. Beyond reproduction, we propose a saliency-guided regularisation extension that introduces an Energy Pointing Game loss, directly supervising the spatial alignment of Simple Gradient saliency maps with target objects. We evaluate our extension across a range of saliency-loss weights and show that explicit saliency supervision substantially improves localisation metrics at a modest cost: the reduction in zero-shot accuracy and adversarial robustness is small at the lowest saliency-loss weight ($w_s=0.2$) and moderate at the strongest setting ($w_s=0.6$). We restructure the codebase for configurable, systematic experimentation. Our reproduction code is available at https://github.com/Andrei3223/Re-Boosting-CLIP-Interpretability.
Chat is not available.
Successful Page Load