What Should Predict-then-Optimize Predict? Feature Design from a Scaffold Distribution
Abstract
In predict-then-optimize (PtO), a machine learning (ML) model predicts an uncertain quantity from context, and a classical optimizer then chooses a decision based on that prediction, inheriting the predictive power of ML and the efficiency and guaranteed feasibility of classical solvers. In PtO's best-known linear form, it suffices to predict the uncertain parameters' conditional mean. We study the nonlinear case, where the mean is insufficient. Existing methods then estimate the full conditional distribution or expected objective, which often contains far more information than the optimization needs. We propose instead designing a few transforms ("features") of the uncertain parameter, predicting the conditional expectations of those features, and passing those to the optimizer. We derive a design principle for the features from a scaffold distribution: a context-free guess of the conditional law, such as pooled historical outcomes. Near the scaffold, the optimal features among all continuous transforms of the same output dimension are Hessian-scaled gradient eigenmodes from one finite eigendecomposition. A small inventory example shows that the designed features capture threshold probabilities that demand means miss, and how distance from the scaffold limits the design.