Risk-Sensitive Deep Optimal Stopping
Abstract
Risk-sensitive optimal stopping arises wherever a single bad outcome dominates a sequential decision, such as option exercise, sequential medical treatment, and safety-critical control. The natural objective is the conditional value-at-risk (CVaR) of the stopped reward. We develop the first theoretical framework and algorithm for risk-sensitive deep optimal stopping. Together they resolve all four structural problems of risk-sensitive RL: state augmentation, score-function pathology, blindness to success, and bootstrapping circularity. We isolate three structural properties of optimal stopping: single-point reward, action-independent dynamics, and sum-zero gradient. They yield Markov optimality on the original state space and uniform boundedness of the per-step pathwise gradient at the deterministic-policy boundary, resolving the first two. From these, we derive DIOS (Distribution-Informed Optimal Stopping), a single-time-scale CVaR estimator on the Rockafellar-Uryasev envelope that resolves the remaining two. A soft penalty replaces the hard tail indicator, and a tail-level annealing schedule initializes training in the risk-neutral regime, where no trajectory is gated out. The inner quantile admits a closed-form update without an auxiliary critic. DIOS admits a finite-time best-iterate stationarity guarantee under standard CVaR non-degeneracy, the first explicit rate for CVaR optimal stopping. On Bermudan max-call options, DIOS achieves the highest mean CVaR and the lowest wall time against five representative CVaR baselines.