A Multitask Protein Language Model For Protein Property and Protein-Protein Interactions
Abstract
Protein language models (PLMs) are typically pre-trained on single sequences, yet many downstream applications require understanding protein-protein interactions (PPIs). We present ProPRINT, a 3B-parameter T5 encoder-decoder model jointly pre-trained on single-sequence inpainting and a novel pair inpainting objective that reconstructs masked residues of one protein conditioned on its interacting partner, using PPI data from StringDB (experimentally validated interactions) and SynthPPI (74.6M synthetic interacting domain pairs derived from the AlphaFold Protein Structure Database). For SynthPPI, we introduce a contact-aware masking strategy that preferentially masks binding-interface residues, integrating structural information into sequence-only models. ProPRINT-3B outperforms ESM2 models up to 3B and MINT on two of five PPI benchmarks and four of eight single-sequence benchmarks, showing that PPI-augmented pre-training enhances representation learning without sacrificing single-sequence representation quality.