Twenty-four minutes of warning: what a reported lead time licenses, and what reporting standards do not ask
Abstract
A median alert lead time is a capability claim: a system reported to warn a median of 24 minutes before an event invites the reading that it will usually give about that much warning. That is not what the number measures on its own. A first threshold crossing can only occur inside the pre-onset window over which a patient was observed. The statistic is therefore set jointly by the model and by the cohort's distribution of those windows, and a single median cannot separate the two. That makes it a reporting problem before it is a modeling one. That a lead time cannot exceed the window it is searched in is definitional; what has not been measured is how strongly the reported first-crossing lead moves with realized, patient-varying available-window length in realistic evaluation pipelines. Across three perioperative cohorts and five model families, at thresholds transported unchanged from a development validation split, median lead among detected event cases rises from the shortest to the longest realized pre-onset window band in all fifteen model-by-cohort combinations. The rise is 26 to 54 minutes. Read beside its operating point, the longest pooled median lead here, 24 minutes --- a median over the event cases a model detected --- stands beside a case-level specificity of 0.406 measured over all control cases; the minimum specificity across the fifteen combinations is 0.379. Two natural responses leave the dependence in place; a horizon-specific detection rate reduces it without removing it, and restricting to a constant window removes it but reveals that the longest median co-occurs with the lowest detection rate. In our author-completed TRIPOD+AI item-level review, the closest requests are items 20a and 20b, which ask for follow-up time as a cohort descriptor; we did not identify an item that asks for the distribution of the window the alert rule actually searches, for a window-conditional timing measure, or for a lead-time summary to be interpreted jointly with its realized operating-point metrics. We set out four candidate disclosures for continuously monitored models, task-specific complements to TRIPOD+AI. A median lead remains a descriptive interval among detected cases, but on its own it supports neither a model-intrinsic nor a cohort-comparable reading of warning capability. We analyze what a reporting convention makes inferable, not what individual readers infer, and do not measure comprehension here.