Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints
Abstract
Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time. Each arrival must receive a decision immediately, and the long-run usage of every resource must stay within its capacity. We study how to learn such a policy from logged observational data. Recent work proposes a decision-blind pipeline that fits one outcome model per treatment by regression, prices each capacitated resource from those models, and assigns each arrival the treatment with the largest predicted outcome minus price \citep{tang2024learning}. We instead train the outcome models end-to-end, differentiating an off-policy estimate of the deployed policy's value through the dual prices themselves. We study two formulations: an exact nonconvex one, and a convex relaxation whose optimum always satisfies the capacity constraints in expectation and whose suboptimality can be explicitly bounded. Every method is evaluated in a queueing simulation to capture the online decision structure. Across six datasets, our two end-to-end variants lead the delay-adjusted policy value wherever capacity is scarce, once a waiting period carries even a small cost. In particular, decision-blind regression predicts ground truth more accurately, but when resource capacities are stringent these baselines frequently violate them and incur much longer queueing delays. Thus, end-to-end training is best suited to settings where resources are genuinely scarce and feasibility matters.