WPWM: When Can a Wearable Patient World Model Tell the Measurement from the Participant?
Abstract
When does a pattern in data reflect the world, and when does it reflect the way we measure it? WPWM tests this distinction through wearable prediction, controlled remeasurement, and intervention inference. Participant-separated models produce a direction-specific activity reconstruction gain and reduce multinight sleep Brier loss from 0.0613 to 0.0552. Exact finite-record bounds preserve the sleep advantage under every binary completion of unknown labels on aligned nights. A fitted, observation-corrected readout generates intervention contrasts before fresh outcomes are supplied. In a 32-case controlled study, component-refitted intervals cover 292/292 issued contrasts in one unbiased-reference condition but 205/295 in its biased counterpart. Moving the same two missing observations to the response-bearing window reduces coverage from 154/295 to 31/292. A randomized sleep analysis and deterministic workflow controls supply complementary evidence. On the 74 evaluation people at landmark horizon 𝐻=14, encouragement contrasts across three frozen checkpoints agree to within 0.02 h (spread 0.016 h;minimal-𝛿 landmark-to-ITT compatibility at 0.014–0.030 h) while incentive contrasts span 0.10 h(0.096 h; minimal-𝛿 0.133–0.229 h): encouragement is model-invariant, incentive is model-dependent, and this divergence surfaces only under intervention queries. The results separate scoped predictive robustness from measurement-related inference failures. They do not establish biological identification or a causal chain from prediction gains to harm. Generated-readout inference is conditional on a frozen neural background. The agentic layer is stated as six design targets with declared evidence states, not as a report of language-model performance.