Cross-Lingual Transfer of Activation-Space Refusal Anchoring Against Adaptive GCG Attacks
Abstract
We study whether a training-free activation-space refusal defense, Latent-Space Refusal (LSR) Anchoring, transfers across languages without multilingual training. Anchoring clamps residual-stream activations toward a refusal direction extracted solely from English data at inference time, with no weight updates. We evaluate the defense on Mistral-7B-Instruct-v0.1 in a single-turn, white-box adaptive setting in which an attacker has access to the defense and re-optimizes against it. In English, the anchor reduces attack success rate (ASR) by 40% for static suffixes and yields a 20% lower ASR than the undefended baseline under adaptive optimization, while approximately doubling the attacker's optimization loss. We additionally conduct a pilot evaluation in Yoruba, Igbo, Hausa, Swahili, Arabic, and Igala. Results are heterogeneous, including zero observed static successes for Igbo and lower adaptive than static ASR for Arabic. Because the multilingual evaluation lacks per-language undefended baselines and uses three prompts per language, these results are descriptive pilot observations rather than isolated estimates of defense transfer. Code, results, and anchor-extraction scripts, but not attack suffixes or harmful completions, are available at https://anonymous.4open.science/r/lsr-anchoring-gcg-defense-F04E/.