Between Serology and Storm: Antiphospholipid Syndrome as a Canonical Stress Test for Patient World Model Reliability
Abstract
Computational work on antiphospholipid syndrome (APS) now spans two well-developed layers: molecular stratification of who is biologically at risk, and diagnostic classification of which disease is present. A third layer — the clinical trajectory layer — asks a different question: for a patient of known serological risk, in a specific clinical context, under a specific anticoagulation decision, what happens next and when? Recurrence prediction in APS is an active field — validated risk scores, a prospective multi-model validation cohort, machine-learning and digital-twin treatment comparisons, and a 500-patient multicentre anticoagulation cohort — but every one of them estimates multi-year recurrence risk conditioned on stable covariates. They do not condition on the transient second hits — hospitalisation, surgery, infection, anticoagulation interruption — that convert latent serological risk into an acute event. This paper argues APS is a canonical stress test for this layer, and makes the argument falsifiable rather than rhetorical. First, it specifies the trajectory-layer estimand and shows, with a worked two-process construction, that when trigger exposure is never time-linked to serology the long-horizon marginal risk stays identified while the short-horizon trigger-conditional treatment contrast does not: two processes matched to an identical 26.2 % 26.2% annual risk differ threefold in post-trigger risk difference. Second, it tests whether the failure is instead a reasoning failure, using a pre-registered 120-trial vignette study across four hypothesised failure modes, three presentation styles and two model tiers; all 120 trials passed, a null result that relocates the failure upstream to data fragmentation. Third, it audits four minimum evidence conditions against PubMed and Europe PMC: anticoagulation exposure is RCT-grade, stratified event rates are partial, and second-hit exposure linkage and diagnostic delay as a modelled covariate remain absent even in the cohorts that do record longitudinal recurrence and anticoagulation intensity, the latter returning zero indexed hits. Fourth, it releases an executable second-hit uncertainty regime that suppresses point-timing estimates and discloses which evidence conditions support each output, audited over 4,800 enumerated states and eight logical invariants — two of which encode gate specifications that invert their own intent. The pattern generalises beyond APS to any disease combining latent risk with a transient trigger.