From Interpretation to Discovery: Sparse Autoencoders as Tools for Plant Genomics
Abstract
Our understanding of genomes is built progressively by years of hard work as well as progress in sequencing technologies. Yet even well-studied plant genomes still contain substantial functional information that remains undiscovered or incompletely annotated. Genomic language models have already shown to capture functional information from purely unsupervised training. We show that the interpretability techniques such as sparse auto encoders (SAEs) not only open the blackbox of such models, but can also work as an instrument themselves. The association between a SAE feature and a biological concept behind it often requires much lower annotated examples than training a classical classifier, and gives deeper insights and transferability to unseen data. First, by examining individual features of plant genomic language model we show the ability to identify likely discrepancies of current splicing annotations. Secondly, we show that alternative splicing may be predicted using a 2-level decision tree with single-feature thresholds in the nodes. The findings generalize to species where no annotated data was used.