Long Context Is Not Enough: An Empirical Case for Dynamic Context Management and Reasoning in ICU Clinical Decision-Support Agents
Abstract
ICU-based AI prediction models generally process a fixed window of routinely measured clinical data, like vital sign-measurements, and compute an output prediction. Recent agentic AI systems leverage reasoning, memory or multi-agent collaboration. However, there is limited work investigating the impact of long context on system performance with prolonged length of stay. We characterize this problem on a large-scale publicly available ICU dataset (MIMIC-IV): ((i) serialized clinical events scale near-linearly with length of stay, with the longest stays exceeding the context window of most deployed LLMs; (ii) token mass is heavy-tailed and dispersed across 1,300+ chartevents variables, so a fixed window cannot ingest a stay and a curated subset necessarily discards most of it; and (iii) feeding an LLM more raw context zero-shot does not improve in-hospital mortality prediction. Naive long-context ingestion is therefore not sufficient on its own, which motivates active context management for reliable ICU agents. Our code is available at https://anonymous.4open.science/r/lcm_icu-325A/