SelfGuard: Self-Supervised Deviation Modeling for Multi-Modal Jailbreak Detection
Abstract
Multi-modal large language models (MLLMs) achieve strong vision–language reasoning ability but remain vulnerable to jailbreak attacks that exploit subtle cross-modal cues to bypass safety mechanisms. Detecting such attacks is difficult because harmful data are scarce and rapidly evolving, while detectors trained on known patterns often fail to generalize. In this paper, we cast multi-modal jailbreak detection as self-supervised deviation modeling and learn attack-agnostic signals from benign data only. We propose SelfGuard, a self-supervised framework, which models benign multi-modal regularities and learns to score structured violations via controlled pseudo-harmful deviations. SelfGuard synthesizes pseudo-harmful samples via controlled transformations that target three signature cues, cross-modal semantic inconsistency, format manipulation, and semantic toxicity. It then learns complementary deviation statistics through multi-task self-supervised learning, including misalignment discrimination, reconstruction-based toxicity modeling, and contrastive format modeling. At inference, SelfGuard performs multi-view deviation estimation by aggregating task-specific deviation scores to identify departures from benign multi-modal regularities. Extensive experiments on various benchmarks show the effectiveness of our method.