Population Structure Invalidates Naive Evaluation of Genome-Based Phage Selection
Abstract
Predictive models built for treatment selection have to be evaluated under the shifts they will encounter in a true setting. I wanted to study a preclinical analogue of predicting Staphylococcus aureus response to bacteriophages from the bacterial genome. The reconstructed benchmark I used contains 253 isolates, 17 reported clonal complexes (CCs), eight phages, and 2,024 quantitative OD600 outcomes. Applying an identical phenotype-blind protein pangenome and fixed ridge pipeline, random cross-validation was conducted resulting in macro Spearman 0.457 and R2=0.132, whereas complete-CC holdout reports 0.190 and -0.237. Random splitting was found to overstate Spearman by 0.267 (CC-bootstrap 95% CI 0.130-0.438) and R2 by 0.369 (0.144-0.694). CC was confirmed as a strong dependence unit without using outcomes due to findings of whole genome similarity, and a separate 45-isolate cohort reproduces random-split optimism (AUROC difference 0.147, 95% CI 0.048-0.254). No prespecified receptor or defense marker passed both multiplicity control and independent-lineage replication. The model was found to be unsuitable in the end for clinical phage selection. The result displays a validation framework for high-stakes treatment-response models: specify the deployment shift, hold out the biological dependence unit, cluster uncertainty accordingly, separate mechanistic hypotheses from validated effects, and require prospective assay-compatible evidence before decision analysis.