Taxonomic Tokenisation and Self-Supervised Pretraining for Transferable Microbiome Representations
Abstract
Taxonomic abundance profiles provide a compact view of microbial communities, but heterogeneous and partially resolved taxonomies make learning transferable representations challenging. We investigate whether taxonomic tokenisation and self-supervised pretraining can address this problem. We assemble Atlas, a corpus of 539,308 microbiome profiles from diverse ecological contexts, and introduce a fallback tokenisation scheme that retains the most specific available taxonomic assignment when genus-level labels are unavailable. Using Atlas, we pretrain the Waypoint family of models, GPT-style causal language models trained to understand microbiomes. We evaluate their learned representations on Compass, a benchmark spanning eight environmental and gut-microbiome prediction tasks. We show that pretraining consistently improves downstream performance and enables favourable scaling with model capacity. Further, models pretrained without human-associated or gut samples still perform well on human gut tasks, demonstrating transfer learning across ecological domains. We find that incomplete vocabulary coverage across independently processed datasets can limit downstream performance, identifying taxonomic tokenisation as an important remaining limitation. Together, these results show that self-supervised pretraining can learn transferable microbiome representations while motivating more robust tokenisation methods for microbiome data.