What Are We Actually Predicting? A Leakage Audit of Clinical Trial Outcome Prediction
Abstract
Clinical trials are costly and frequently fail to reach completion because of futility, safety concerns, recruitment challenges, or other operational constraints. Predicting discontinuation risk early could help sponsors prioritize portfolio review and identify potentially vulnerable trials before substantial resources are committed. Machine learning approaches have increasingly been applied to clinical-trial registry data for this purpose, using both structured trial characteristics and unstructured textual information. However, a fundamental question remained: were these models learning prospective signals of discontinuation, or were they exploiting information that became available only after a trial had already begun to fail? We investigated this question using the Aggregate Analysis of ClinicalTrials.gov (AACT), focusing on the prediction of trial discontinuation from registry text and structured metadata. The cohort comprised 154,421 interventional trials with a terminal recorded status, of which 18.5% were discontinued. The central challenge was that AACT represented the current state of a registry record rather than a historical snapshot of what was known at trial registration. We identified three leakage pathways. First, recorded actual enrollment could encode information about the eventual outcome, whereas only anticipated enrollment would have been available prospectively. Second, registry summaries and eligibility criteria could be revised after discontinuation, introducing post-hoc language referring to termination, suspension, slow accrual, or futility. Third, the same registry export also contained a field recording why a study had stopped (whystopped), populated almost exclusively for discontinued trials, so its presence alone encoded the label. We excluded it by construction, but any pipeline drawing features from the registry record as a whole rather than from named fields would have inherited it. In addition, because completed trials constituted the majority class, conventional accuracy and F1 could provide an overly optimistic assessment of performance. To address these issues, we developed a leakage-controlled evaluation protocol that restricted enrollment features to anticipated values, redacted post-hoc outcome-related vocabulary while preserving cohort size, and evaluated models using a chronological train/test split. We combined ClinicalBERT representations of trial summaries and eligibility criteria with structured trial metadata and compared the resulting model against metadata-only, TF-IDF, text-only, and simple reference models. We evaluated performance primarily using discontinuation-class PR-AUC and recall at fixed alert budgets, reflecting the practical task of prioritizing a limited number of trials for review. Our results demonstrated that evaluation design substantially changed the apparent predictive performance. Under a controlled random split, the full model achieved a PR-AUC of 0.337 ± 0.005 and ROC-AUC of 0.703 ± 0.006. Permitting actual enrollment increased ROC-AUC to 0.851 ± 0.004, an apparent gain of +0.148 attributable to leakage rather than improved prospective prediction. In contrast, removing the text leakage screen had essentially no effect (Δ ROC-AUC = −0.001). The whystopped field was starker still: a single test for a non-empty value identified discontinued trials with a precision of 1.000 at a recall of 0.880, with no model and no training at all. Under a leakage-controlled temporal split, the model achieved a PR-AUC of 0.468 ± 0.001 and ROC-AUC of 0.736 ± 0.000. Within a review budget of 20% of the test partition (n = 30,884), the model surfaced 38.6% of all discontinued trials. These findings showed that what information was allowed into the prediction problem could matter more than the choice of model itself. More broadly, our work highlighted that reliable clinical machine learning required defining the prediction moment before optimizing the predictor. For longitudinal registry data, explicitly accounting for information availability, temporal generalization, class imbalance, and operational decision constraints was essential for distinguishing genuine predictive signal from retrospective artifacts. We argued that this principle extended beyond clinical-trial prediction: before asking how well a model predicts, researchers should first ask what information the model was actually allowed to use to make that prediction.