When Similar Models Disagree: Validation-Induced Decision Instability in Clinical Trial Portfolio Selection
Abstract
Pharmaceutical development is plagued by high costs and attrition rates, and leveraging AI to predict clinical outcomes is an active area of research. This study examined how validation and model-selection choices altered fixed-capacity Phase 2 portfolios. Candidate prediction models were developed using point-in-time data from 522 industry-sponsored, single-agent trials: 116 human-annotated Clinical Trial Outcome (CTO) labels and 406 model-derived labels. Several validation splits were implemented, including random, temporal, asset-grouped, and temporal-with-asset-purge validation. Within each validation path, an ML pipeline was selected using AUROC, Brier score, or standardized expected net present value (eNPV) as the objective; its predictions were then used to form a 20-program portfolio. Eleven selected ML pipelines were evaluated on 237 human-annotated outcomes. Future AUROC ranged from 0.708 to 0.754, while top-20 Jaccard overlap ranged from 0.053 to 1.000. Holding the selection objective fixed, validation paths produced 7–9, 6–9, and 6–9 successes out of 20 across the three objectives. The largest realized gap was three successes, equivalent to $1.159B under common benchmark assumptions, although pairwise intervals included no difference. The results demonstrate the outcome instability that can result across commonly used validation methods for AI-based decision making.