What Makes Multimodal Black-Box Attacks Succeed? Distributional Clues from Image and Text Candidates
Abstract
Black-box attacks on image–text matching rely on limited target-model feedback to search for effective image and text modifications. The relationship between paired image and text representations offers an additional source of guidance, motivating us to investigate its connection to attack performance. We introduce a distance-guided evolutionary search framework that combines target matching scores with paired image–text distances from an independent encoder. The framework generates candidate pairs through image and text variation and combines target-model feedback with paired-distance guidance to select candidates according to a preference for larger or smaller paired distances. The selected population carries this preference into subsequent search. Experiments on BLIP-base ITM with MS COCO show that the two guidance directions move the mean paired distance in opposite directions. Across CLIP and SigLIP2 guidance encoders, selection toward smaller distances achieves lower average scores and comparable or higher attack success rates than neutral and random selection. These results suggest that paired image–text relationships provide useful information for guiding multimodal black-box attacks.