Reinforcement Learning Agents Are Swimmers
Abstract
To date, the field of reinforcement learning (RL) has primarily focused on developing discounted methods, which are designed to optimize the short-term or transient performance of RL agents. Conversely, the field has been reluctant to explore and invest in methods that are designed to optimize the long-term or steady-state performance of RL agents. In this position paper, we challenge this reluctance under the framing that RL agents should be viewed as competitive swimmers. That is, we argue that, like competitive swimmers, RL agents cannot hope to win all races, or solve all tasks, based on transient performance alone. To this end, we first motivate the need for RL agents that can adequately optimize both the transient and steady-state performance. Then, we summarize recent works which show that discounted methods have deficiencies when it comes to optimizing the steady-state performance. Finally, we argue in favor of more attention and investment from the community towards average-reward methods, which are designed to optimize the steady-state performance. Namely, we challenge common misconceptions associated with average-reward methods, and highlight recent advancements which suggest that such methods present a viable and potentially superior alternative.