Model-Robust and Sequentially Valid Moment Certificates for Chance-Constrained Stopping
Abstract
We study a chance-constrained stopping problem in a finite Markov decision process (MDP) with continue/stop actions, where cumulative cost must remain within a given budget with high probability. When the true transition model is unknown, a learned simulator or digital twin provides an estimated model for comparing stopping policies that determine which states to stop in and which states to continue from. For each policy, we compute its expected reward and the mean and variance of cumulative cost through linear systems, and use Cantelli's inequality to determine whether the chance constraint is satisfied without tail rollouts. We further introduce a robust margin that accounts for transition-model error and an online update rule that never delays stopping, thereby preserving the chance guarantee after replanning. Experiments compare the method with sampling-based screens and illustrate the nominal reward--safety tradeoff, the effect of model error, and the benefit of nested online tightening.