Module Specialization in Transformer Factual Recall under Correlated Fact Distributions
Abstract
The standard mechanistic interpretation of transformer factual recall attributes storage to deep MLP modules and routing to attention. We study how module roles change when the fact distribution carries subject, relation, or answer-collision structure rather than independent facts. Which module performs the recall depends on the correlation axis and on the model's capacity load. Probes and matched-counterfactual patching disagree systematically on every motif: probes credit the deepest MLP while patching credits the first attention module. We show that probes and patching measure structurally different quantities---cumulative linear decodability versus irreplaceable causal contribution---which under correlated fact distributions need not coincide. A parameter-counting argument singles out subject-axis collision as the unique compression that the attention-routing pathway can exploit cheaply, predicting an order-of-magnitude crossover for that pathway class that we observe empirically: the layer-2 probe gap collapses sharply for subject collisions, and only subject-axis specialization survives the capacity-tight regime. Edge-level path patching uncovers a three-step chain (Attn -> MLP1 -> MLP2) whose dominance tracks the predicted threshold, with a causal-scrubbing test corroborating that the chain is causally relevant on every motif and sharpest in the subject regime. A dense joint sweep shows that subject and answer compressions compete rather than amplify, with the dominant axis set by the stronger compression. On natural CounterFact strata across four frozen pretrained LMs (Pythia-1.4B/2.8B, LLaMA-3.2-1B-Inst/3.1-8B), edge-level path patching reproduces the synthetic ranking: the answer-collision stratum is the most localized on every model. A two-hop extension shows the binding-locality structure carries to compositional facts.