Does Your Trial Simulator Know What the Drug Does? An Intervention-Blind Test of Interventional Validity on 8,011 Registered Trials
Abstract
Patient world models are increasingly proposed as simulators of clinical trials, and recent surveys organise their capabilities as a ladder that rises from temporal prediction, through action-conditioned prediction, to counterfactual rollouts that can support decisions. Position on that ladder is currently assigned by reading papers rather than by measurement. We turn the counterfactual rung into a measured quantity. From 30,119 completed phase 2/3 trials with results posted to ClinicalTrials.gov we construct 11,274 within-trial contrasts (treatment-arm outcome minus comparator-arm outcome on the primary endpoint) from 8,011 trials, and we score a simulator not on how well it predicts a held-out arm but on how much of its accuracy depends on knowing which intervention was tested. The instrument is an intervention-blind reference: the same model, trained on the same trials, with every intervention-identifying field withheld. Interventional skill is the error the full model removes relative to that reference. On a temporal split (results posted before/after 2023; 1,020 test contrasts from 740 trials) a registry-trained gradient-boosting simulator removes a quarter of the error of a trivial baseline, but its interventional skill is -0.05 (95% CI -0.13 to 0.02, five seeds), and the null replicates on continuous endpoints, on alternative temporal cuts, in placebo- and active-controlled strata, across five alternative architectures, and when the intervention is retrieved from earlier trials of the same molecule, while the sampling-noise ceiling of the task is 0.91. Language models reading the same protocol fields behave differently: a frontier model removes 0.62 of baseline error with the intervention shown and 0.39 with it hidden, an interventional skill of 0.38 [0.27, 0.49] that is stable across repeated runs, reappears on continuous endpoints (0.28) and in a smaller model of the same family (0.18), and rises to 0.69 [0.45, 0.82] on the 76 contrasts from 50 trials whose primary completion postdates the model's training cutoff, ruling out memorisation of posted results. Held-out accuracy alone ranks the two systems in the same order but for the wrong reason; only the blind control separates learning an endpoint's scale from learning what a treatment does. We add a reliability layer (conformal intervals whose coverage slips to 0.83-0.86 under temporal shift, and selective prediction that halves error at 25% answer rate) and provide the benchmark, code and prompts as anonymised supplementary material.