One Label, Several Events: Measurement Failures in the Standard Clinical-Trial Outcome Benchmark
Abstract
Machine-learning models of clinical-trial success are trained and compared almost entirely on one public resource: the Trial Outcome Prediction (TOP) benchmark derived from ClinicalTrials.gov. We link all 11,601 of its trials back to their registry records and audit what the benchmark actually measures. Three problems appear, in the label, the features, and the strata. First, the binary outcome label pools distinct events. Between 5.2% and 46.9% of trials, depending on phase, have no efficacy primary endpoint at all and are judged on safety or pharmacokinetics, and those trials are labelled successful far more often: 67.1% against 49.1% in Phase II (adjusted OR 1.82, p = 2.2 × 10⁻⁸), with the same direction in Phase I and Phase III. Second, two post-hoc fields leak. Registry status, shipped beside the label, inflates ROC-AUC by +0.136 to +0.262, which exceeds the total reported gain of every architecture published on this benchmark; and the registry's own enrollment and site counts, realised after the fact for 91% of records, independently add +0.076 to +0.237. Third, the feature and stratum channels are corrupted: 54.1% of distinct SMILES strings are shared by more than one drug name, 81.5% of trials carry a structure shared by more than five names, one structure is assigned to 377 arm names of which 98.9% are placebo arms, and 6 of 16 reference drugs carry a pharmacologically adjacent compound's formula. Using central nervous system trials as a worked example, we then show the therapeutic-area stratum is itself incoherent: Phase III success ranges 2.66× across CNS subareas (χ² = 64.8, permutation p ≤ 5 × 10⁻⁵) while pooled CNS is statistically indistinguishable from pooled non-CNS at every phase, and a fifth of CNS-labelled trials are not CNS by the registry's own MeSH hierarchy. One reassuring result: the temporal split is genuine, which we verified rather than assumed. The remedy is better reporting rather than better architectures.