Electro-Prot: Learning Protein Representations from Electrostatics
Paula Sharoubeem ⋅ Sumeet Kothare ⋅ Jacky Chen ⋅ David Koes
Abstract
Protein Language Models (PLMs) trained on sequence, as well as structure-aware variants that integrate structural information, have emerged as powerful tools to learn protein representations applicable to a wide range of downstream tasks. However, neither of these frameworks explicitly accounts for the electrostatic interactions that govern protein functions. We investigate whether a continuous electrostatic field can serve as a pretraining target to encode general protein representations. We present Electro-Prot, a 6.7 million-parameter encoder-decoder transformer model trained to predict Poisson-Boltzmann electrostatic potentials at arbitrary points around a protein. Coordinate queries cross-attend to atoms to predict the electrostatic potential as a continuous field. Electro-Prot attains $r=0.97$ (MAE = 0.6 kT/e) on a held-out test set. On downstream tasks, our embeddings match or outperform PLMs in interaction-centric tasks despite utilizing a dataset that is orders of magnitude smaller. PLMs remain stronger on GO, binding site detection and protein family classification. Combining Electro-Prot and PLM embeddings improves results in structure similarity, protein-protein interface, Gene Ontology, and binding site detection, suggesting that Electro-Prot learns a complementary signal not captured in evolutionary statistics. Our results suggest that electrostatic pretraining provides a viable, parameter- and data- efficient signal for learning representations relevant to interaction-centric tasks.
Chat is not available.
Successful Page Load