The Assistant Persona Differs Across Languages in Activation Space
Abstract
Large language models behave differently across languages, and are noticeably less safe in lower-resource ones. This gap is well established behaviorally, but not representationally. The directions used to describe model internals, including the assistant persona, are almost always extracted from English and assumed language-agnostic. We extract the assistant persona direction independently in each of 31 languages in Qwen3-32B, and find that the resulting axes do not point in a common direction. Their cosine similarity to the English axis ranges from 0.378 to 0.911. Moreover, each language's assistant persona direction and its magnitude are correlated with the language's resource level. The same holds for the default assistant state, which lies further from the helpful assistant pole in lower-resource languages than in higher-resource languages. Finally, steering against a language's own axis amplifies persona-based jailbreak more than steering against the English axis does. These results indicate that the model does not carry a single assistant representation across languages. More broadly, the assistant persona may need to be an explicit target of multilingual alignment.