Perturb-LM: Strong Baselines Change the Interpretation of Biomedical Language-to-Morphology Retrieval
Abstract
Machine learning models are increasingly proposed as interfaces for complex biomedical datasets, yet apparent improvements can be difficult to interpret when learned systems are not evaluated against strong non-neural controls or under conditions designed to reduce metadata shortcuts. This distinction is particularly important when model performance may be communicated as evidence of biological understanding. We developed Perturb-LM, a leakage-aware benchmark for evaluating whether biomedical language representations support natural-language retrieval of Cell Painting morphology profiles. Perturb-LM uses 4,524 quality-controlled JUMP CPJUMP1 profiles represented by 904 morphology features, with 4,190 profiles containing non-missing perturbation labels included in retrieval evaluation. Queries were constructed from descriptive biological metadata while excluding treatment and sample identifiers, target sequences, plate, well, batch, and other direct lookup features. We compared identifier-stripped TF-IDF with frozen BiomedBERT embeddings and a train-only ridge projection from BiomedBERT representations into morphology space. Evaluation included predefined held-out-plate and held-out-treatment settings with additional plate-and-well retrieval exclusions. In the primary held-out-plate condition excluding same-plate and same-well matches, 180 of 1,079 queries were evaluable. Identifier-stripped TF-IDF achieved mean average precision (mAP) 0.1574 (5,000-bootstrap 95% CI, 0.1367–0.1795), compared with 0.0092 for unaligned BiomedBERT and 0.0332 for projected BiomedBERT. Although projection improved the frozen representation, projected BiomedBERT underperformed TF-IDF by ΔmAP −0.1243 (paired 95% CI, −0.1519 to −0.0946). The same direction was observed across all four predefined evaluation settings. A targeted leakage audit detected no direct identifier or lookup-key overlap within held-out-treatment queries under the defined audit procedure. These results show how the interpretation of biomedical ML performance can change when learned representations are compared with strong, transparent controls. Improvement over an unaligned neural representation did not constitute evidence of successful language-to-morphology transfer. Perturb-LM therefore provides a case study for more cautious biomedical ML capability claims and supports explicit baseline, leakage, held-out, and uncertainty-aware evaluation when communicating model performance.