Same Values, Different Clock: Separating Measurement Timing from Physiology in ICU Risk Explanations
Abstract
Two ICU patients can have identical laboratory results and receive different risk scores, because they were measured on different schedules. Staff measure the patients they are already worried about, so when a patient was measured is a clue to how sick they are, and a model can learn it. When the model then explains an alert by naming "lactate," the clinician hears a fact about the patient's lactate; the model may partly mean that somebody kept ordering lactate. This paper gives the first construction that tells those apart inside a single record, holding the number of draws fixed so that when a patient was measured is separated from how often, and the first measurement of how far apart they are. For each of 5,424 held-out ICU patients in the PhysioNet/Computing in Cardiology Challenge 2019 data, we rebuild the record twice: once keeping every measured value, its order, and the number of measurements while moving only the hours at which they were taken, and once keeping the hours while redrawing the values. The schedule's share of the resulting explanation movement is 0.39 for a gradient-boosted tree and 0.55 of its risk-score movement. A recurrent model that receives no schedule feature gives 0.55 and 0.62, two unrelated explanation methods agree, a tree with every schedule input removed gives exactly 0.000, and the effect reappears on a second dataset at a fixed hour-24 landmark with a mortality outcome. Because that ratio depends on how hard the value arm pushes, we run the value arm at six strengths: the share moves from 0.67 to 0.39, the preregistered arm is the floor of that range, and when the two arms displace the model's input equally the share is 0.51.