Beyond Drift Detection: Deployment Monitoring for Selective Prediction
Abstract
Selective prediction lets a model defer uncertain cases rather than act autonomously, which is particularly useful in settings such as assistive activity recognition. Its decision threshold is typically calibrated before deployment, but sensor failures, changes in user behavior, or other distribution shifts can make that threshold unsuitable. A drift alarm only indicates that the data have changed; it does not tell us whether the deployed threshold remains feasible, another threshold could restore operation, or autonomy should be suspended. We formulate deployment monitoring as an authorization problem over a finite set of candidate thresholds. At each checkpoint, the controller evaluates selective risk, autonomous coverage, and deferral-targeting value, which measures whether deferrals target cases where intervention is useful, for both the deployed threshold and its candidate replacements. It then returns one of five actions: Continue, Retune, Watch, Escalate/Relabel, or Hard Revoke. On synthetic data and the OPPORTUNITY activity-recognition benchmark, similar distribution shifts lead to different reference states: Continue, Retune, or Suspend. On PAMAP2, using disjoint monitoring and verification data, risk-only retuning produces 54 observed infeasible authorizations across 17 of 300 perturbation conditions, whereas the proposed controller and its guarded-retuning variant produce zero observed infeasible authorizations in 1,200 evaluated windows. Deployment monitoring should therefore determine whether selective autonomy remains justified, not merely whether the data distribution has changed.