CareQueue: Reinforcement Learning Methods for Intelligent Patient Prioritization
Ayush Jain ⋅ Julia Kryłowicz ⋅ Beloslava Malakova ⋅ Anusha Asthana
Abstract
Delays in intensive care can have serious consequences, particularly for patients with time-sensitive conditions, such as sepsis. This work investigates whether reinforcement learning (RL) methods trained on retrospective ICU data (MIMIC-IV dataset) can support intelligent patient prioritization. We compare the following models: Behavior Cloning (BC), Batch-Constrained Q-Learning (BCQ), Implicit Q-Learning (IQL), and Double Deep Q-Learning (DDQN), then evaluate their priority scores in a queue simulator and compare them with first-in-first-out (FIFO). The learned policies reduce median waiting time from 41.83 minutes under FIFO to approximately 27--28 minutes, while mean waiting times remain similar. For high-severity patients (SOFA $\geq 6$), median waiting time decreased from 38.88 minutes under FIFO to 24.30, 25.42, and 24.51 minutes under BCQ, IQL, and DDQN, respectively. These results show substantial improvements in waiting times for high-severity patients (SOFA $\geq 6$), suggesting that model-based prioritization can help direct limited treatment capacity toward patients with increased clinical severity, although at the cost of increased tail waiting times. Lastly, the current evaluation relies on a simplified simulated queue with synthetic patient arrivals and service times, and further validation under more realistic ICU conditions is needed before drawing conclusions about the clinical benefit.
Chat is not available.
Successful Page Load