Characterizing Latent-Space Adversarial Training for LLM Safety
Maria Smirnova ⋅ Yarin Gal ⋅ Lin Li
Abstract
Latent-space adversarial training for LLM jailbreak robustness optimizes adversarial perturbations within constrained regions of learned representation spaces. Enlarging that region gives the training adversary access to a larger perturbation set, so training with a larger radius is often expected to improve robustness. We test this expectation through Continuous Adversarial Training (CAT), where input-embedding perturbations are constrained within an $\ell_2$ ball of radius $\epsilon$. We conduct controlled CAT radius sweeps across three model--recipe configurations spanning Zephyr and Llama 3, evaluated against four jailbreak attacks together with task capability and benign behavioral utility. We show that larger radii improve robustness against some attacks over parts of the evaluated range rather than providing a consistent ordering, and the benefit depends on the attack and model--recipe configuration. We further find that benign compliance can deteriorate substantially at radii where general task capability remains comparatively stable, and that some larger-budget settings incur additional utility loss without additional measured robustness. We then characterize what enlarging the perturbation region does to training itself. We find that perturbations remain closest to their starting token embeddings and several larger settings show substantial between-seed variability. Finally, we show that targeted over-refusal supervision can learn the targeted behavior without consistently recovering broader benign compliance, while degrading robustness. Together, these results show why perturbation budgets in learned representation spaces must be interpreted through their realized perturbations and robustness--utility effects rather than treating configured radius as a direct measure of adversarial-training strength.
Chat is not available.
Successful Page Load