Conservative Offline RL for ALS Trajectory Emulation: Falsification Tests and the Limits of Intervention-Aware Patient World Models
Abstract
Patient world models intended for clinical-trial simulation must emulate longitudinal disease trajectories while avoiding unsupported extrapolation under partial observability and limited historical support for intervention regimes. We study this problem in amyotrophic lateral sclerosis (ALS) using a POMDP-inspired conservative offline reinforcement learning (CQL) trajectory emulator trained on PRO-ACT data. On the PRO-ACT DREAM release, CQL improves short-horizon ALSFRS-R recovery relative to behavior cloning and yields more stable long-horizon rollouts than the unregularized Q-learning configuration; its strongest relative performance occurs on an outcome-defined rapid-progressor stratum, while a linear population-slope baseline remains stronger for slow, near-linear trajectories at the longest evaluated horizon. A configuration sweep shows every conservative setting holds seed variance low, though the unregularized comparator also differs in rollout rule and so does not isolate the penalty. We further identify a reliability failure mode: with a one-dimensional observed state, apparent perturbation robustness can arise from input-insensitive policy collapse; augmenting the observed state with forced vital capacity restores nonzero input-dependent rollout responses while retaining lower perturbation sensitivity than behavior cloning. Active-arm and riluzole analyses on the full PRO-ACT registry do not provide affirmative validation of intervention-effect counterfactuals, owing to weak or absent treatment signal, treatment-selection bias, and a design that manipulates a covariate rather than a learned action. These results establish both the promise and the limits of conservative offline RL for ALS patient-world modeling. We consolidate them into an explicit validation matrix that separates what our evidence supports from what it does not, and that names the evidence still required before trial-simulation use: explicit intervention semantics, target-trial design, support diagnostics, calibrated uncertainty, and evaluation in informative treatment settings.