StyREx: Training-Free Red-Teaming of MLLMs via Structured Style Recipe Search
Abstract
Visual presentation can alter the safety behavior of multimodal large language models (MLLMs), motivating style-based visual red-teaming. Existing training-based style optimization, however, faces two limitations. First, it couples repeated target feedback with target-specific optimization of a large image editor, incurring substantial training and weight-maintenance overhead. Second, holistic style labels entangle visual factors, preventing evidence sharing across styles and offering little guidance after a failed trial. We introduce StyREx (Style Recipe Exploration), a training-free framework that casts style enhancement as structured, feedback-guided recipe search. StyREx decomposes holistic styles into eight atomic operations and searches over singleton and pairwise recipes. A frozen image editor executes each recipe, while continuous feedback from the target MLLM and safety judge updates a global-to-category empirical Bayes posterior. No-repeat Thompson sampling balances exploitation and exploration under a bounded query budget. All neural models remain frozen, and target adaptation resides only in a lightweight statistical state. Experimental results demonstrate that our StyREx outperforms its training-based counterpart in both attack success rate (ASR) and harmfulness score (HS), while eliminating backpropagation and target-specific parameter optimization.