AdaptiveCareArena: Decision-Facing Evaluation of Patient World Models Under Healthcare Resource Constraints
Abstract
Patient world models are typically evaluated by predictive fidelity. However, their clinical value depends on the decisions those beliefs support under the finite sensing, human attention, and intervention capacity real care systems have. We introduce AdaptiveCareArena, a real-trajectory replay benchmark evaluating patient models by the resource-constrained decisions they support. We instantiate two resource axes on independent cohorts (CrossCheck and CES). First, we show need-aware review allocation robustly reduces missed high-severity events across budgets. Second, for measurement selection, we find an adaptive policy offers little practical advantage over a strong fixed short-form, despite an oracle confirming exploitable headroom exists. Finally, we identify observation-action entanglement, a validity failure where benchmarking conflates observation availability with human review. AdaptiveCareArena treats observation, review, and intervention as distinct events to avoid this, exposing the model as a substitutable interface.