A Cross-Lingual Dissociation Between Representational Geometry and Causal Mediation
Abstract
Refusal in language models is mediated by a single residual-stream direction, and a direction fitted in one language can suppress refusal in others, a transfer attributed to the parallelism of these vectors in representation space. Because parallelism and causal transfer co-occur in that setting, parallelism has not been tested independently as evidence of shared mechanism. We test this relationship using a role-evidence conflict construct in Qwen3-8B. Rank-1 directions fitted independently in English, Chinese, Spanish, and Hindi on mutually disjoint scenarios converge at signed cosine 0.76–0.90 against a shuffled-label null of 0.26–0.42, with alignment emerging abruptly at layer 20 and peaking at 0.919 at layer 24. However, ablating these directions at five depths across the convergence plateau leaves the role-conditioned decision gap within the label-permutation null in nineteen of twenty language-by-depth cells; the sole exception also exhibits the largest specificity violation. The same intervention code reduces refusal on held-out prompts by 0.844. We further identify format-driven near-determinism, option-order sensitivity, and failures of assigned-referent identification despite near-ceiling task comprehension. These results show that cross-lingual representational convergence does not, by itself, imply shared causal mediation.