Necessary but Not Sufficient: Causal Dissociation and Cross-Model Variability in a Chess Transformer's Checkmate Circuit
Suriya D Saravanakumar ⋅ Om Lala ⋅ Pranav Ramesh ⋅ Laksh Patel
Abstract
We identify two attention heads, L1H1 and L2H7, that are each independently necessary for a chess-playing transformer to detect a forced mate in one move: zero-ablating either head collapses accuracy from a 51.0\% baseline to 2.5\% and 18.0\% respectively. Isolating these two heads together with a third candidate, L2H6, while ablating the remaining 61 heads drives accuracy to 0.0\%, so the candidate set fails a stringent sufficiency test despite each head's individual necessity. The two necessary heads also differ sharply in internal structure: L1H1's output-value (OV) circuit concentrates 84.7\% of its squared Frobenius energy in a single singular direction, compared with 10.7\% for L2H7. Where measured, both heads' necessity generalizes across mate-depth and mating-piece splits. Replicating the ablation on three further checkpoints trained on different data distributions shows that which head dominates is not fixed: L1H1 dominates on two Lichess-trained models, L2H7 on two Stockfish-influenced models, with a single-head dissociation in one case (SF\_MIX: ablating L2H7 gives $d=-1.31$, $p=6.7\times10^{-40}$, while ablating L1H1 gives $d=-0.002$, $p=.395$). Together these results argue that identifying a component whose removal breaks a behavior is not enough, by itself, to establish a stable, self-contained explanation of that behavior.
Chat is not available.
Successful Page Load