Hindsight Is 20/20: Making Regret Meaningful in Decision-Focused Learning
Abstract
Benchmarking decision-focused learning (DFL) requires three choices: an oracle, a normalizer, and an aggregation rule. Auditing 24 papers, we find there is no consensus. Most authors report regret against an oracle with perfect foresight of realized costs. Worse, most inferences drawn from such regret metrics are tainted: they inflate regret on instances that are arguably easy'', make flat performance appear to trend, compress differences between methods, and can even reverse rankings. We illustrate each effect on published studies. Our diagnosis is confined to linear contextual problems, and our audit is a sample, not a census. Finally, we offer best practices. For synthetic experiments, we recommend a new metric, \emph{SAA-Relative Regret}, which avoids these pitfalls by using aninfinite data, infinite compute'' oracle and normalizing by the regret of the non-contextual SAA policy. On real data that oracle's cost cannot be computed; we propose an estimator built from a champion policy, prove it asymptotically efficient once the champion's own regret is negligible, and show that estimating any policy's regret is equivalent to estimating the optimal policy's performance. We close with open questions.