What Transformer FFNs Never See: Theory, Diagnosis, and Lightweight Remediation
Tinghe Zhang ⋅ Yucheng Xiao ⋅ Alex Lamb
Abstract
In Transformer attention, distinct weight vectors $\boldsymbol{\alpha} \neq \boldsymbol{\alpha}'$ can produce the same aggregated representation $\mathbf{V}\boldsymbol{\alpha} = \mathbf{V}\boldsymbol{\alpha}'$. When $\operatorname{rank}(\mathbf{V}) \leq n-2$, configurations with entirely different dominant source tokens collide on a set of positive Lebesgue measure, a condition satisfied in over 97% of attention heads in BERT-Base. Such collisions are harmful: the downstream FFN receives identical inputs despite the attention having attended to entirely different source tokens, and must therefore produce identical outputs regardless of which tokens actually dominated, fundamentally limiting any computation that requires sensitivity to attention source, such as multi-hop reasoning, attribution, or knowledge retrieval. No increase in depth, width, or data can compensate, because the routing weights are discarded before the FFN is reached. The routing weights $\{\alpha_{ij}\}$, however, are available within the same forward pass and can be retained as a compact side-channel to restore the missing signal. We call this structural information loss *routing non-identifiability*. In BERT-Base, value-matrix rank deficiency suppresses $k \geq 107$ routing dimensions (over 84% of degrees of freedom) before any FFN computation. To quantify the practical gap, we introduce the Routing Reconstruction Task (RRT): standard Transformers achieve 31.4% source-identification accuracy while oracle routing statistics reach 94.6%, a 63-point gap consistent with an information bottleneck in the FFN input. We then propose the Route-Aware FFN (RA-FFN), a drop-in module that appends four compact routing statistics to each FFN block at under 1.1% additional parameter cost, recovering over 75% of the oracle gain. RA-FFN improves all benchmarks tested: at 110M scale, gains cover multi-hop reasoning, language modelling, and translation, with multi-hop gains $6\times$ larger than single-hop; at 7B scale on Llama-2-7B and Mistral-7B, gains extend further to mathematics (GSM8K) and general reasoning (MMLU, ARC-Challenge, HellaSwag), ordered by routing sensitivity.
Chat is not available.
Successful Page Load