Emergent, Context-Dependent Group Representations in Genomic Foundation Models
Jake Kovalic ⋅ Rachel Zhang ⋅ Siavash Golkar ⋅ Shirley Ho
Abstract
Interpreting geometric structure in learned representations is difficult when the structure of the underlying variable is itself unknown. Biology provides rare cases where that structure is specified independently of the model. During translation, coding sequence is read in three-nucleotide codons, assigning each coding position one of three reading-frame phases, while outside coding regions this phase has no translational meaning. This provides a controlled setting in which to ask whether pretrained genomic models recover a known, context-dependent structure, where it emerges, and whether it is functionally used. Across four genomic foundation models, codon phase is detectable in hidden representations. In MIMIC, a multimodal foundation model trained across RNA and protein modalities, this signal is concentrated in a low-dimensional subspace that forms a near-equilateral geometry consistent with the cyclic group $C_3$. We causally validate this interpretation through targeted latent interventions: rotating this subspace by $120^\circ$ shifts the model’s reading frame for protein decoding by one nucleotide, while the inverse rotation induces the opposite shift. Decoded frame shifts under successive rotations follow $C_3\cong\mathbb Z/3\mathbb Z$, and latent rotations further compose predictably with sequence-level insertion and deletion operations. Together, these results show that a biologically specified cyclic structure can be recovered as an explicit, causally active geometry in the latent space of a pretrained model.
Chat is not available.
Successful Page Load