Benchmarking under Sparse Ground Truth: Cross-Modal Retrieval for Drug Discovery from Microbial Natural Products
Abstract
Many of medicine’s most powerful drugs are synthesized by microbes, yet the therapeutic potential of microbial chemistry remains largely uncharacterized (Demain, 2014). Researchers study this chemical potential through two complementary data modalities. Genome sequencing identifies biosynthetic gene clusters (BGCs) that encode what a microbe may produce, while untargeted mass spectrometry records fragmentation spectra of the molecules it actually produces. Linking BGCs to spectra could turn genome sequencing into a search tool for unknown molecules (Schorn et al., 2021). However, experimentally verifying a single link can require months of laboratory work. We formulate this problem as bidirectional cross-modal retrieval: ranking candidate spectra for a BGC and candidate BGCs for a spectrum. Because discovery requires predictions for organisms absent from training, the task is out-of-distribution by design. Evaluation must therefore determine whether a model learns transferable biosynthetic patterns or memorizes associations with organism provenance. To support this evaluation, we curated available BGC--spectrum links into evidence tiers using public paired-omics and experimentally validated BGC resources (Schorn et al., 2021; Medema et al., 2015). Approximately 3,000 links are supported only by chemical-structure matching, while 46 publicly accessible links have direct experimental confirmation and form the complete high-confidence test corpus. We reproduced and audited state-of-the-art systems for linking genomic and metabolomic data (Leão et al., 2022; Eldjárn et al., 2021). Although their published scores were reproducible, several systems could not score most of the verified links. Their pipelines rejected the input or excluded links under internal constraints. Two tools returned no predictions for any uncharacterized BGCs in the validation genomes. We also found evaluation choices that can inflate performance: reporting accuracy at ranks where chance retrieves 60\% of matches, removing low-confidence queries instead of counting them as failures, and using validation sets as small as 14 examples. Several methods are also unsuitable for prospective discovery because they represent BGCs through metadata about previously observed organisms or rely on unsupported pipelines. To address these limitations, our evaluation framework tests models only on organisms absent from the training data. It compares each model with a matched chance baseline and a randomized control that preserves BGC observation frequency. It also reports hit rate alongside abstention and counts every query, including those for which a method returns no answer. We are evaluating existing systems and our model on all 46 verified links. Preliminary results suggest that as reference coverage decreases, methods based on organism metadata approach the randomized baseline and abstain more often. Performance above both controls would indicate transferable biological signal. This work contributes an evidence-tiered dataset, a reproducible audit and failure-case analysis of prior systems, and a leakage-controlled benchmark for sparse biological supervision. The audit identifies a central design requirement: BGC representations should encode gene content and avoid reliance on organism provenance, consistent with recent sequence-based approaches (Patin et al., 2025). More broadly, the benchmark connects data-centric machine learning, robustness to distribution shift, and out-of-distribution generalization in an extreme low-label regime. Rigorous evaluation is essential because current methods can fail silently on uncharacterized organisms, which may contain biosynthetic pathways that yield new medicines.