Controlled Substitution Reveals Where Learned Forecasting Components Fail
Abstract
Learned forecasting components can improve one diagnostic while concealing failure on another. Controlled substitution exposes two such failure modes. Under price censoring, a raw volatility update silently under-predicts tail-event probabilities; substituting the conditional tail moment repairs probability level, whereas refitting the same corrected model contributes almost nothing further. In multivariate scenario generation, a learned diffusion residual degrades the marginal and joint structure already supplied by matched empirical donors. The diagnostics disagree across objectives: probability level improves without a reliable ranking gain, distributional scores favor the empirical generator while pooled calibration is unresolved, and the learned residual increases rather than collapses forecast spread. Treating each component and its retained reference as a two-point search therefore reveals failures hidden by system-level evaluation and identifies which objective each failure affects.