On objective mismatch in molecular retrieval from tandem mass spectrometry
Abstract
Identifying molecular structures from LC-MS/MS spectra is a central problem in computational metabolomics, often approached by predicting molecular fingerprints and retrieving candidates via similarity search. Despite widespread use, the relationship between training objectives for fingerprint prediction and downstream retrieval performance remains poorly understood. In this work, we show that these objectives are fundamentally misaligned. Adopting a decision-theoretic perspective, we derive novel regret bounds that characterize when Bayes-optimal predictors for fingerprint similarity diverge from those for molecular retrieval. Our analysis reveals that optimizing standard similarity-based losses can provably degrade retrieval performance, and that the extent of this mismatch depends on the similarity structure of candidate sets. Empirically, we validate our theory on the MassSpecGym benchmark, demonstrating a Pareto frontier between fingerprint accuracy and retrieval metrics across commonly used loss functions. These results expose an inherent trade-off in fingerprint-based molecular identification and provide principled guidance for the design of learning objectives in computational mass spectrometry.