An architecture for a human digital twin from longitudinal data at healthcare system scale
Abstract
A human digital twin (HDT) is a model of an individual patient that is personalised to their data, updated as new observations arrive, and predictive of their future clinical states. At longitudinal, healthcare-system scale such a system could forecast how a patient develops under alternative treatments before one is chosen. Here we set out a proposed architecture for building a HDT, and the evidence standard it would have to meet. Recent advances in electronic health record (EHR) foundation models still fall short of the bidirectionality paradigm that distinguishes a twin from a digital shadow, and we locate that shortfall in three gaps: no persistent patient-specific state, no causal identification for the counterfactuals they generate, and erosion of the outcome signal once the model is deployed. We propose three co-designed layers addressing each in turn: an intervention-conditional population model exposing a queryable latent state, an agentic personalisation layer holding per-patient memory, and a causal-validity layer that pre-specifies estimands and predictimands. Robust evaluation of a HDT, we argue, will likely require "silent-mode" assessment followed by a cluster-randomised trials with clinical endpoints. The HDT that earns the name will be the first to show all three layers working together on real patients, evaluated to the rigour a randomised trial imposes.