TrialCalibre1: An Agentic System for Trial Emulation, Benchmarking, and Calibration
Abstract
Healthcare databases can help evaluate treatment effectiveness for questions both addressed and left unanswered by randomized trials. Doing so requires linking clinical evidence, study design, and analyses that can be run on the data. We present TrialCalibre1, an agentic system supporting reference-trial replication, direct target-trial emulation, and BenchExCal workflows that link benchmarking, extension, and calibration sensitivity analysis. We evaluated an ASCET-based benchmark and an extension comparing clopidogrel with aspirin among adults aged 81–90 years using synthetic PCORnet-style claims. Two reviewers assessed benchmark extraction, protocol decisions, and implementation. After discussion, 53/62 and 50/62 assessable SQL checks were rated met; both reviewers rated 12/13 assessable R checks met. Protocol and proxy disagreements remained. A separate ten-question routing development set yielded 10/10 route matches and 7/7 applicable reference-trial matches. The agent benchmark and extension included 6,069 and 1,602 patients. The extension risk ratio was 0.848 (95% CI 0.569–1.184), with a calibrated sensitivity estimate of 1.274 (0.439–3.700). The human implementation provided descriptive comparisons, not same-data validation. Dose proxies, unpopulated covariates, residual imbalance, implementation deviations, and ignored cross-cohort covariance limited interpretation. This case demonstrates a connected workflow whose protocols, code, and results can be reviewed.