Recommendations as Designed Experiments: Active Inverse Optimization for Identifiable Preference Learning under Hard Constraints
Farzin Ahmadi ⋅ Kimia Ghobadi
Abstract
Inverse optimization (IO) personalizes constrained recommendations, for instance, diets under clinical guidelines, by recovering the objective that renders observed decisions optimal, then re-optimizing. It has been shown that passive observation does *not* necessarily identify the objective, but instead recovers a polyhedral cone of parameters and more data from the same behavior cannot shrink it. However, a recommender system can be viewed as a designed experiment that needs to offer targeted interventions that are informative while maintaining or improving user's baseline goals (e.g., dietary goals). We formalize this as active, intervention-aware, online IO where accept/reject responses cut a tracked admissible set of preferences; slates composed within an $\varepsilon$-indifference family of guaranteed-improving options constructively satisfy the excitation condition that passive data violate; and a randomization-identified exposure term corrects the bias the system injects into its own data. We validate this system using a closed-loop simulation that shows 100% non-worsening recommendations while reducing intervention-biased preference error from 0.20 to 0.13 by round 60.
Chat is not available.
Successful Page Load