MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Junyoung Park ⋅ Namgyu Park ⋅ Sechan Lee ⋅ Yoon-Chan Jhi ⋅ Jihoon Cho ⋅ Sangdon Park
Abstract
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important scenario for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, but existing turn-level credit signals for jailbreaking are often too coarse for multi-turn interaction. We propose a unified turn-level credit assignment framework for Group Relative Policy Optimization (GRPO) in multi-turn jailbreak learning. Our method assigns group-relative learning signals directly at each turn by decomposing immediate and future credits. We call this approach **decomposed credit GRPO (DC-GRPO)**. In contrast, existing methods often suffer from credit misassignment, for example by applying a single trajectory-level score to the entire dialogue. Across multiple victim LLMs and benchmarks, our dynamic-weighted DC-GRPO achieves an average $ASR@3$ of $98.26\%$ for 5-turn jailbreaking, outperforming existing state-of-the-art methods such as SEMA, which achieves $86.58\%$, and TROJail, which achieves $86.23\%$. These results highlight turn-level group-relative credit assignment as a simple and effective ingredient for scalable multi-turn automated red teaming.
Chat is not available.
Successful Page Load