Evaluating LLM-Based Clinical Digital Twins for Longitudinal Patient Forecasting
Mohammad Amin Kamaleddin ⋅ Xiaoyang Liu ⋅ Nghia Le
Abstract
Trustworthy generative artificial intelligence (AI) for health requires more than high predictive accuracy: a person-conditioned model should demonstrably depend on the correct individual's longitudinal state, and any gain from retrieval should not be misattributed to language model reasoning. We introduce a leakage-controlled evaluation framework for large language models (LLMs) that separates correct-person profile alignment from population-level predictability and separately tests the LLM's incremental value beyond a fixed personalized retrieval prior. Using the Medical Expenditure Panel Survey (MEPS) longitudinal public-use file, we froze 27 future outcomes spanning physical function, mental health, and healthcare utilization; sampled 250 participants per domain across five participant-isolated folds; and applied the same protocol to three model runs. A one-to-one shuffled-person control preserves realistic profiles while breaking profile—outcome alignment. Across nine model—domain replications, correct-person profiles improved accuracy in eight after Holm correction: by 14.1—19.5 percentage points (pp) for physical function, 5.4—12.3 pp for mental health, and 16.6—18.2 pp for utilization in two models, while one utilization replication was null. In contrast, an LLM given the identical personalized 12-neighbor distribution improved over deterministic k-nearest-neighbor (kNN) argmax in only three of nine replications after correction. No reference LLM condition significantly outperformed target-wise longitudinal persistence after correction, and a prespecified $2\times2\times2$ prompt factorial found no corrected main effects of evidence selection, population priors, or prediction memory. Person-aligned history therefore carries reproducible forecasting signal, but that signal should not automatically be attributed to the LLM or interpreted as mechanistic twin fidelity. Person-swap and matched-prior controls provide practical audit tests for longitudinal and retrieval-augmented generative models in health.
Chat is not available.
Successful Page Load