Repeated-Measure Leakage, Distribution Shift, and Reliability under Partial Observation in Patient World Models
Abstract
Patient world models are increasingly proposed for longitudinal prediction, intervention-aware reasoning, and clinical-trial simulation, but intervention validity is downstream of a more basic requirement: the underlying predictive state must generalize across patients, survive realistic shifts and missing observations, and expose failure through meaningful reliability signals. We evaluate these prerequisites in a deliberately narrow, falsifiable setting: short-horizon digital-biomarker forecasting from PhysioNet GaitPDB. A file-level audit yields 165 usable participants, 306 recordings, and 51,129 context–future pairs. Using persistence, ridge, MLP, GRU, Transformer, and a compact JEPA-style latent predictor, we build an evaluation ladder that progressively removes raw temporal overlap, same-recording familiarity, and same-patient familiarity before testing unseen-patient generalization. For GRU, NMSE rises from 0.1227 under random-window splitting to 0.1393 after eliminating all raw train–test overlap and to 0.1961 under patient holdout. Among 54 participants with repeated recordings, exposure to a different recording from the same patient improves GRU NMSE from 0.2177 to 0.1556 (Holm p = 1.9 × 10⁻⁶), while a recording-excluded identity hypothesis is not supported at the participant level, preventing a causal memorization claim. Under valid patient holdout, MLP and Transformer are statistically indistinguishable (0.1769 vs. 0.1766; Holm p = 0.291). Study shift, a four-times-longer prediction gap, and partial observation further degrade performance; under 50% temporal masking, Transformer NMSE rises to 0.611 while MC-dropout predictive variance falls. We do not claim a longitudinal or intervention-aware simulator. Instead, the results support a prerequisite evaluation stack—patient separation, repeated-measure controls, shift, missingness, and uncertainty validation—that should be passed before stronger patient-world-model claims are trusted.