Emergent Codon Frame Representations and Causal Translation Steering in Genomic Foundation Models
Abstract
Scientific interpretability requires evidence that a decoded representation is causally faithful to the underlying scientific variable. We use translation reading frame, whose structure and downstream consequence are independently known, as a controlled test case in genomic foundation models. Across four models, coding sequence representations exhibit stronger period-3 structure than non-coding sequence representations, and codon phase is detectable with linear probes of hidden states. In MIMIC, a multimodal foundation model jointly trained across RNA and protein modalities, the leading label-free representation of coding RNA reorganizes from nucleotide identity toward codon phase with depth. We fit a two-dimensional codon-phase subspace and show that manipulating it causally steers the reading frame used for RNA-to-protein translation. Rotating this subspace by (120^\circ) shifts the model’s reading frame for protein decoding by one nucleotide, while the inverse rotation induces the opposite shift. Repeated rotations and sequence edits compose according to the expected modulo-three structure, while matched controls do not induce signed frame shifts. These results provide a controlled example of validating scientific representations through causal intervention against independently known biological structure.