Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
Abstract
We propose a test-time defense mechanism against adversarial attacks: Unlike existing methods that rely on feature filtering or smoothing, which can lead to information loss, we propose to ``combat noise with noise'' by leveraging stochastic resonance to enhance robustness while minimizing information loss. Our approach introduces small translational perturbations to the input image, aligns the transformed feature embeddings, and aggregates them before mapping back to the original reference image. This can be expressed in a closed-form formula, which can be deployed on diverse existing network architectures without introducing additional network modules or fine-tuning for specific attack types. The resulting method is entirely training-free, architecture-agnostic, and attack-agnostic. Empirically, the method achieves state-of-the-art robustness on image classification and provides the first generic test-time defense for dense prediction tasks, including stereo matching, optical flow, and monocular depth estimation. Across adversarial attacks, it recovers up to 68.1% of the accuracy loss on image classification, 71.9% on stereo matching (MAE), 29.2% on optical flow (EPE), 68.3% on depth estimation (SqRel), and 83.8% on vision-language alignment (cosine similarity).