Preventing Model Extraction in Contextual Linear Optimization
Abstract
We investigate model extraction in contextual linear optimization, where a deployed predict-then-optimize system first uses a prediction model that maps a context to a cost vector and then returns an optimal solution to a downstream linear program. Although the predicted cost is hidden and many cost vectors can induce the same solution, we show that observing solutions can still lead to prediction model extraction by an attacker. For linear prediction models, we introduce a loss-calibration condition under which global inverse-feasibility consistency implies recovery of the prediction model. This condition reveals that an attacker can conduct model extraction by leveraging inverse optimization with boundary sampling, i.e., sample queries that result in cost vectors that are close to the boundary where the optimal solution changes. To prevent and defend from such attacks, we show the decision-maker should perturb decisions based on cost vectors that are near boundaries, where the attacker information gain is high but objective regret is small. We prove prediction-recovery and regret trade-off guarantees under the aforementioned assumptions. We also evaluate the attack and defense strategies on synthetic knapsack instances and a short-video recommendation task. Our results show that contextual optimization systems can be vulnerable to model extraction, and that targeted decision perturbation can substantially increase extraction error while preserving decision quality.