MedGenesis: Auditable World-State Updates for Clinical Research Agents
Abstract
Lifelong research agents must preserve not only facts, but also the evolving state of claims, cohorts, tools, uncertainty, and safety constraints across extended workflows. Clinical research is a demanding testbed because a stale cohort definition, endpoint, or statistical assumption can corrupt later actions. We present MedGenesis, a clinical research agent built around auditable world-state updates. Its memory binds each claim to a clinical question, cohort, endpoint, evidence source, tool output, uncertainty, failure mode, and safety boundary. At each round, MedGenesis proposes falsifiable hypotheses, selects an executable action by expected information gain, uncertainty reduction, and a safety prior, executes modular research skills, and writes back only critic-reviewed state updates. We evaluate the system on ClinicalResBench, a 1,697-question clinical-research benchmark, ClinicalRepBench, 40 paper-reproduction tasks, and agent-behavior tests for hallucination, falsifiability, action ranking, and information-gain calibration. In matched-budget evaluations, MedGenesis outperformed evaluated frontier language-model, biomedical-agent, and code-executing baselines, while remaining short of complete study reproduction, where expert clinical and statistical review is still required. Three clinical traces illustrate the intended use: bounded evidence synthesis, trial-enrichment hypothesis generation, and mixed-evidence treatment monitoring. The results position clinical research as a high-stakes domain for studying lifelong agents whose memory, tool use, and alignment must remain inspectable over time.