Learning Risk-Averse Policies with Two-Stage Deep Decision Rules
Abstract
The Two-Stage Deep Decision Rules (TS-DDR) framework combines a neural policy with deterministic recourse for multistage stochastic optimization. We introduce Risk-Averse TS-DDR, which applies conditional value-at-risk (CVaR) to cumulative economic cost and penalizes infeasible policy outputs in expectation. The different trajectory weights require separate derivatives from the recourse multipliers. We show when this separation is exact, derive the policy update, and prove convergence of the exponential moving average threshold while the policy changes. Two experiments compare the proposed method with stochastic dual dynamic programming (SDDP) using a full-horizon CVaR formulation, one on portfolio optimization and a second on hydropower planning over a full year of operations. The results show policy quality comparable to or better than SDDP, with substantially lower policy computation time.