Online Decision-Focused Learning under Semi-Bandit Feedback
Abstract
Decision-focused learning (DFL) improves on the Predict-then-optimize (PtO) pipeline by training predictive models to minimize decision regret rather than prediction loss. Most DFL methods assume offline full-feedback data with complete action cost vectors. Many real-world decision systems instead operate in an online setting: in each round, the agent acts using past feedback and observes costs only for the selected components, yielding semi-bandit feedback. Bandit methods optimize regret in this regime, but train models with predictive updates rather than decision-focused updates. In this paper, we formalize online semi-bandit DFL: decision-focused learning for online optimization problems with semi-bandit feedback. The most straightforward extension would impute unobserved components with the model's own predictions and compute the decision-focused gradient. But this creates a self-confirming imputation degeneracy: imputation errors can flip counterfactual decisions that define the decision-focused gradient, biasing the representation update and reinforcing errors in later imputations. To address this degeneracy, we propose BayesianSPO, which combines an analytically updated Bayesian last-layer head with masked decision-focused learning. The Bayesian head supplies both (i) uncertainty to balance exploration and exploitation and (ii) stable imputations for unobserved components, while masking unobserved component gradients prevents those imputations from directly corrupting the representation update. Across knapsack, shortest-path, and portfolio benchmarks, BayesianSPO achieves the lowest mean cumulative regret on all five instances and the improvement is statistically significant on four of them.