Surrogate Calibration for Transferable Adversarial Attacks against Black-Box MLLMs
Abstract
Transfer-based attacks expose serious adversarial vulnerabilities in closed-source Multimodal Large Language Models (MLLMs) by crafting adversarial examples on off-the-shelf surrogate models. However, this paradigm often drives adversarial examples toward surrogate-specific local optima, where they deceive only the surrogate model but fail to transfer to black-box MLLMs. Although existing methods mitigate this issue at the attack level, they leave the surrogate itself unchanged. To address this gap, we propose Surrogate Calibration (SCal), a novel model-level refinement that enhances downstream transfer-based adversarial attacks in a plug-and-play manner. Specifically, SCal employs a feature-preserving loss to maintain the surrogate's representational utility, while a distributional regularizer encourages its gradient field to stay closer to the natural data manifold. With calibrated surrogates, downstream attacks generate update directions that exhibit superior generalization across black-box MLLMs. Extensive experiments demonstrate that SCal boosts current attacks to new heights across various surrogates. For instance, it elevates baseline M-Attack's ASR on GPT-5.4 Mini to 53.5\%, substantially surpassing the state-of-the-art MPCAttack (29.7\%). These results expose critical vulnerabilities that must be addressed for trustworthy multimodal systems. Code is at: https://anonymous.4open.science/r/SCal_code-B0BC