First-Token Broadcasters: Mechanistic Analysis of Language Identity and Distributed Robustness in Transformers
Arjun Pillai ⋅ Anjelo J Laroza ⋅ Christian Hoang
Abstract
How is language identity represented and causally maintained in transformers, and to what extent do the mechanisms identified in one language set transfer to other languages and training regimes? We introduce Language Identity Head Ablation (LIHA), a causal intervention that zeros each attention head individually and measures the resulting language switch rate. Across our multilingual evaluation sets, LIHA identifies heads whose ablation has disproportionate effects on output language. In GPT-2, a small set of first-token broadcaster heads---led by L6H1 (switch rate 0.32, 3.23$\sigma$ above the population mean)---attend persistently to the first prompt token across generation steps. The combination of persistent first-token attention and causal sensitivity suggests that these heads participate in maintaining output language. Following ablation, we observe statistically significant redistribution of first-token attention ($p < 10^{-5}$) toward heads in higher layers for the ablations tested; this pattern is consistent with downstream compensation, although attention redistribution alone does not establish a feedforward causal cascade. We also compare Qwen2.5-1.5B-Base and Qwen2.5-1.5B-Instruct, which are matched in architecture and parameter count but differ in training history. The base model is nearly flat (max SR=0.016, 200/336 heads at SR=0.0), whereas the instruct model concentrates causal influence at layer~0, led by L0H5 (SR$=$0.224, 8.93$\sigma$ above mean). Finally, GPT-2 heads identified from the European-language experiments show no observed switch effects on the tested Chinese and Russian prompts, while other early-layer heads do show effects. Because script, typology, and tokenization vary together in this comparison, these results do not isolate a specifically script-driven mechanism.
Chat is not available.
Successful Page Load