PaxBench: A Multimodal Sequence Benchmark for Protein Abundance Prediction
Abstract
Protein abundance links genotype to cellular state and is useful for comparing genes, organisms, and engineered systems. PaxDb contains broad abundance mea- surements, but its heterogeneous records have not yet been organized into a shared benchmark for testing protein, CDS, codon, and mRNA representations under the same target. The difficulty is that abundance labels must be recovered from diverse PaxDb sources, matched to species-specific sequences at several molecular levels, and evaluated despite incomplete modality coverage, limited metadata, and variable annotation quality. To address this, we introduce PaxBench, a carefully cleaned and aligned seven-species PaxDb-derived benchmark with matched protein, CDS, codon, and mRNA inputs, and use it to test whether biological foundation models benefit from multimodal fusion. On the cluster-aware split, the full multimodal representation reaches Spearman 0.789, compared with 0.740 for the best single representation, while random evaluation reaches 0.834. Cross-species settings also remain predictive. These results suggest that protein, coding-sequence, codon, and transcript representations provide complementary predictive information under this benchmark. PaxBench provides a compact setting for comparing foundation models, feature baselines, adaptation methods, and interpretability analyses for sequence-based abundance prediction.