HyProt: A Long-Context Hybrid Attention-SSM Protein Language Model and a Matched Comparison with Transformers
Abstract
Protein language models are at one of two extremes in how they handle sequence context. Transformers such as ESM-2 attend to every pair of residues, which is precise but costs compute quadratic in sequence length. State-space models such as LC-PLM instead compress the sequence into a fixed-size memory, keeping the cost linear but giving up the exact recall of distant positions. Hybrid attention-SSM models occupy a useful middle ground for text and DNA. In this paper, we introduce HyProt, an 8M-parameter hybrid attention-SSM protein language model pretrained on UniRef50 with a masked language modeling objective and a 1024-token context, and compare it against a transformer trained with the same data, optimizer, and token budget, so that any difference comes from the architecture. HyProt performs on par with the matched transformer on ProteinGym. Evaluated out of distribution, on held-out sequences up to 16x the 1024-token training context, perplexity diverges at roughly 4K residues: below this the transformer is slightly ahead, while HyProt avoids the sharp degradation the transformer suffers above it. The throughput crossover falls around 8K residues. These results position the hybrid backbone as a practical middle point on the precision-efficiency spectrum for protein modeling. This is work in progress; we are scaling model size and adapting training for downstream tasks such as protein-protein interactions.