Utility-Constrained Policy Optimization
Abstract
Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents, however, the framework does not support risk-sensitive constraints. This can be problematic: For example, CMDPs allow for optimal solutions that, in order to satisfy the risk-neutral constraints, mix infrequent catastrophic behaviors and frequent, overly conservative ones. Moreover, empirical results in multiple previous works suggest that enforcing stricter, risk-sensitive constraints can improve agent performance even when measuring it in a risk-neutral way. In this work, we introduce a simple yet powerful methodology for constrained RL, consisting of an extension of CMDPs that removes the limitation of risk-neutral constraints, and an algorithm for solving the resulting constrained problem. As a convenient side-effect of our framework, it is not necessary to fix constraint limits in advance of training the agent, provided that a sensible range is known. This increases policy flexibility and, in practice, allows for adjustments to these limits at no extra training cost. Besides benefiting from the generality of the framework, our agent shows strong performance in practice, consistently matching or outperforming existing baselines in several Safety Gymnasium benchmark tasks.