Risk-Aware Action Repetition via Expected Skip Evaluation
Abstract
Open-loop action repetition—hereafter also referred to as a skip—enables reinforcement learning agents to reduce decision frequency by executing selected actions for adaptive durations. A key challenge is how to evaluate the value of such a skip. Existing Skip-MDP methods typically bootstrap from the greedy value of the terminal state, implicitly assuming that the agent can immediately resume near-optimal control after the skip. This assumption is unreliable during learning, when the underlying action policy is stochastic, inaccurate, and continually changing. It can also amplify optimistic estimation errors, distorting the learned preference over skip lengths. We propose Expected Skip Evaluation, a reformulation of skip-value learning that evaluates skip terminal states through expected future control rather than idealized greedy control. This reflects imperfect control during training while preserving the optimal skip-value fixed-point as control becomes reliable. We instantiate this principle in RARe (Risk-Aware Repetition), a practical framework for discrete and continuous action spaces. Experiments across grid-world, continuous-control, and safety-critical benchmarks show improved sample efficiency and better-calibrated skip decisions.