CohortWorld: Aligning Real-World Data-Driven Diagnostics and Language Models
Ziquan Wei ⋅ Tingting Dan ⋅ Guorong Wu
Abstract
Language models (LMs) can handle complex real-world tasks since they can act as agents in interactive virtual environments with verifiable rewards as trustworthy feedback.
However, real-world data-driven diagnostics has no such environment. Patient-level health records sit behind data use agreements, and the public health environments are synthetic, access-gated, or narrowed to urgent care, resulting in the absence of real-world evidence (RWE) verifying LMs to avoid hallucination in medical LMs.
To bridge this gap, we present CohortWorld, an environment generated from UK Biobank (UKB) in which each state is a real-world cohort keyed on disjoint histories of disease, procedure, and medication.
By adaptive history granularity matching, disjoint $k$-anonymous cohorts ($n$=29,884) with real-world data are gathered as states from $>$500k UKB subjects with different diseases labeled by ICD codes. CohortWorld contains an RWE calibration benchmark and an RWE-diagnostics simulation suite, namely CohortCal and CohortThink, respectively.
When testing both closed-source and open-source LMs on CohortCal, they struggle around/below 0.5 AUROC with bad slope scores on disease risk ranking. In contrast, Qwen3.5-9B can outperform Sonnet 5 using CohortWorld, which implemented RWE matching and risk altering by counterfactual actions. Together, CohortWorld advances the trustworthiness of artificial diagnosis and enables LMs to reason like a human doctor based on verifiable RWE in a releasable and actionable virtual environment.
Chat is not available.
Successful Page Load