Between Life and Death: Examining Failure Modes of Sparse Reward Designs in Healthcare RL
Abstract
In reinforcement learning (RL) for healthcare, reward functions often encode clinical endpoints like survival and death. This results in a sparse reward structure with non-zero rewards only at terminal transitions. However, the exact reward magnitude and combination of survival rewards and death penalties vary across studies, with the implicit assumption that these choices are interchangeable. In this work, we theoretically and empirically examine three common sparse reward designs: survival-only, death-only, and mixed. We prove that, under the assumptions of terminal-only rewards, guaranteed absorption, and no discounting, the corresponding value functions of the three designs have an equivalence relationship and lead to the same optimal policy; while breaking any of these conditions can break the equivalence in consequential ways. On simulated environments, we verify the theoretical results and demonstrate how relaxing these assumptions affect the equivalence relationship. Finally, we consider a healthcare domain of sepsis treatment based on real-world clinical data, in which all theoretical assumptions are violated. Overall, we find that the death-only reward design consistently underperforms and leads to indecisive policies, while survival-only and mixed rewards perform similarly with differences arising from how they handle trap states and interact with intermediate rewards and discounting. Our work provides the first systematic characterization of these failure modes in healthcare RL, offering theoretically grounded guidance for practitioners on reward design choices.