AdaCipher: Open-Ended Cipher Search for Jailbreaking LLM Systems
Abstract
Despite the continuous development of safety mechanisms, which drastically decrease effectiveness of many natural-language attacks, ciphers open a new opportunity for jailbreaks. We study the influence of ciphers on model capabilities, the behaviour of guardrails, and the effectiveness for jailbreaks. Previously explored adaptive cipher-based jailbreaks are limited to combinations or selections of ciphers from predefined sets. Hence, we propose AdaCipher, an automated pipeline that uses adaptive search to discover transformations beyond a predefined set. AdaCipher outperforms all considered predefined ciphers on all four target models without guardrails, achieving up to 79\% ASR. Moreover, AdaCipher-discovered transformations notably improve the effectiveness of the considered natural-language jailbreaks against guardrailed systems. Overall, our results indicate that ciphers represent an important and exploitable dimension of the jailbreak space that should not be overlooked when evaluating the robustness of LLM safety mechanisms.