Mechanistic Discovery in a Genomic Foundation Model Reveals a Donor-Acceptor Computation
Cameron Berger ⋅ Padmaja Mohanty ⋅ SRI BHANU GUNDU ⋅ Jeyashree Krishnan ⋅ Ajay M Rangarajan
Abstract
Genomic foundation models learn long-range biological dependencies, yet their internal computations remain poorly understood. We apply mechanistic interpretability to this setting, using causal interventions to study donor--acceptor pairing in SpliceBERT, a masked language model trained on primary RNA from 72 vertebrates. Corrupting the paired acceptor changes the masked-donor margin by $0.872$ logit units ($95\%$ CI $[0.771,0.969]$), while matched non-partner controls are near zero. Nested interventions resolve four roles: positional geometry supplies an address; early attention heads and feed-forward layers construct an acceptor-local payload; Layer 3 Head 1 combines them into a donor-side write; and later heads deliver this state to the prediction. Yet this causal core is not a sparse autonomous circuit: complement isolation preserves only $0.54\%$ of coupling, and a validation-split faithfulness curve reaches $89.9\%$ only at 64 of 103 nodes. Separately, nucleotide-invariant intervention shows that off-register rotation of a learned period-3 positional component reduces coupling more in coding-sequence blocks than in matched upstream-intron blocks, while a linear-margin chimera assay finds stronger coupling for matched coding-boundary registers. These results provide a mechanistic account of one genomic-model behavior and turn an unresolved model regularity into a bounded, falsifiable hypothesis for external testing.
Chat is not available.
Successful Page Load