StereoPep: Do Molecular Models Understand Stereochemistry? A Benchmark on Synthetic Diastereomeric Peptides
Michael Desgagné ⋅ Amirabbas Kazeminia ⋅ Kübra Kaygisiz ⋅ Bradley Pentelute ⋅ Marinka Zitnik
Abstract
Existing peptide models saturate natural-sequence proteomics benchmarks, but whether they capture the underlying physicochemical dynamics or merely pattern-match on sequence statistics remains unclear. Public peptide datasets offer no axis along which to test this: non-canonical residues and stereochemical variation are essentially absent. We introduce StereoPep, a benchmark of 48,788 chemically synthesized peptides spanning 11 libraries, produced in the laboratory via high-throughput flow synthesis and characterized experimentally by reverse-phase liquid chromatography retention time. Unlike model-generated synthetic data, every peptide in StereoPep is a real molecule with a real chromatographic measurement. The dataset is constructed around matched diastereomeric pairs (sequences identical in primary structure but differing at a single stereocenter), providing a direct test of whether models capture three-dimensional stereochemical information. StereoPep enables progress across multiple fields: it serves as a benchmark for stereochemistry-aware molecular representations and geometric deep learning, supports adapting protein language models to non-canonical residues, and underpins emerging proteomics multiplexing strategies that encode samples through retention-time shifts. Across specialized retention-time models, graph-based architectures, and protein language models, all approaches predict overall retention reasonably well (Pearson $r \approx 0.80$) but collapse on diastereomer pair discrimination (AUC $\approx 0.55$), revealing that current representations encode atomic composition far better than 3D structure.
Chat is not available.
Successful Page Load