Deciding Pipeline Order Under Joint Severity Shift: From Local Diagnostics to Target-Law Performance
Abstract
Operational machine learning models process inputs through multi-stage pipelines whose transformations do not commute. When deployment environments experience distribution shift, the cost-optimal execution sequence can invert, which renders static baseline configurations suboptimal. We formalize pipeline sequencing as a prescriptive decision problem under a target deployment distribution. Operational decisions are governed by the joint loss-plus-cost differential, evaluated through simultaneous confidence intervals with explicit switching tolerances and safe abstention thresholds. We prove that validation means, local curvature, and response attributions answer distinct structural questions, whereas prescriptive optimization requires direct estimation of target-law expected loss. For estimation, we analyze a paired counterfactual estimator for randomized baseline distributions and develop MIXCURV, a nested orthogonal score estimator with cross-fitting that eliminates confounding under covariate-dependent sampling. Across eleven multi-stage pipeline benchmarks, validation diagnostics diverge from deployment performance, optimal sequences reverse across severity shifts, and MIXCURV reliably recovers true interaction signs where unadjusted estimators exhibit sign errors.