Self-Attention-Guided Transferable Gray-Box Adversarial Patch Attacks on Vision Transformers
Abstract
Vision Transformers (ViTs) and their variants have achieved remarkable success across a wide range of computer vision tasks, including safety- and security-critical applications. However, the security implications of exposing self-attention information in ViTs remain largely unexplored. Existing adversarial attacks on ViTs typically rely on white-box assumptions with unrealistic access or black-box settings with limited effectiveness. This paper investigates a realistic gray-box threat model motivated by modern deployment practices, where self-attention information is accessible through interpretability or analysis interfaces. We show that such attention transparency introduces a previously underexplored vulnerability in Vision Transformers. We then propose Targeted Patch Perturbation (TPP), an attention-guided adversarial framework that localizes semantically important patches and restricts perturbations to these regions. The proposed method enables effective adversarial attacks without requiring gradient access to the target model. By explicitly aligning perturbation placement with attention-dominant regions, TPP achieves substantially stronger attack effectiveness than random patch perturbation and existing black-box baselines under identical patch budgets. Unlike existing patch-based approaches, TPP establishes a direct and principled link between attention patterns and patch placement, enabling stable and semantically aligned selection of influential image regions. By exploiting attention as a model-intrinsic indicator of semantic importance, TPP enhances both attack effectiveness and transferability across diverse Vision Transformer architectures and variants. We instantiate the proposed framework in two variants: TPP-W and TPP-G. TPP-W represents a TPP-based white-box formulation that serves as a strong upper-bound baseline for attention-guided patch attacks under full internal access. More importantly, TPP-G corresponds to a TPP based gray-box formulation that leverages attention information derived from the victim model while optimizing adversarial perturbations using surrogate models. This design reflects realistic deployment constraints and enables effective attacks without requiring gradient access to the target model. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrate that the proposed TPP framework is highly effective under both white-box and gray-box settings. In particular, the gray-box variant TPP-G consistently outperforms strong black-box baselines and, under comparable patch budgets, can even surpass white-box random patch strategies while preserving high perceptual quality. Further analysis across diverse Vision Transformer architectures reveals that attention-guided patch localization plays a critical role in enhancing transferability, highlighting a systemic reliability risk associated with attention-based interpretability in transformer-based vision models.