Stochastic Decision Horizons for Survival-Constrained Reinforcement Learning
Abstract
We propose stochastic decision horizons (SDH), a theoretically grounded framework for reinforcement learning with per-step violation signals. Rather than treating violations as expenditure of a cumulative budget, SDH shortens the effective decision horizon when violations occur, reducing both current reward credit and future bootstrapping. We show that this construction corresponds to a survival-chance constrained problem and retains a Bellman-compatible objective. Using Control as Inference, we derive two off-policy, regularized algorithms with different post-violation semantics. Absorbing-state semantics ends the decision process, so only surviving decisions pay policy-information cost, yielding AS-SAC. Virtual termination stops reward credit while keeping the decision process alive, yielding VT-MPO with KL-constrained policy improvement. To connect SDH with standard cumulative-cost CMDPs, we introduce violation-depth profiles. The two objectives agree under a single-scale violation profile and diverge when frequent shallow violations mix with rare deep ones. Experiments validate both the method and this predicted scope. On the 90-muscle H2190 humanoid in Hyfydy, VT-MPO matches the domain-specific state-of-the-art EWA baseline in peak gait realism, with its highest observed gait-match checkpoint at 21M environment steps versus 90M for EWA and substantially more stable training. On Safety Gymnasium, violation-depth profiles correctly predict the regimes in which SDH achieves strong reward-violation trade-offs.