Predictive Representation Alignment Improves Generalization in LLM Safety
Abstract
Jailbreak attacks keep the same unsafe intent while changing the prompt's surface form. This shift is a problem for representation-level defenses: a defense can learn the attacks it sees during training instead of the intent beneath them. Existing defenses reshape hidden states for harmful prompts, but they do not directly use paired clean and adversarial views of the same request. We propose Predictive Representation Alignment (PRA), a paired-view regularizer that uses an auxiliary mapping function to pull adversarial-view hidden states toward clean-view commitment states. PRA can be added to several harmful-side regularizers, which lets us test whether clean-adversarial alignment improves robustness beyond the base defense objective. Across six attack families, PRA lowers average ASR under StrongREJECT in 8 of 11 matched backbone and defense combinations without degrading benign capability. The clearest effect is on probe transfer: final-layer harmfulness probes recover OOD accuracy across model families and scales, from 8B backbones to Qwen3-32B, with gains up to +65 percentage points.