The Critic Is Only a Proxy: Regret Optimization for Reinforcement Learning
Yikai Wang ⋅ Shang Liu ⋅ Guanting Chen ⋅ Jose Blanchet
Abstract
Actor-critic algorithms improve a policy through a learned critic, but the critic is only an informative and imperfect proxy for environmental return. Pessimism to mitigate the proxy issue is therefore intrinsic even without distribution shift. Existing methods primarily optimize pessimistic value; we instead minimize worst-case regret against the action that would be optimal under the same plausible critic. We use a simple example to show that distributionally robust regret optimization (DRRO) is less pessimistic than distributionally robust value optimization (DRO), avoiding over-pessimism while controlling proxy exploitation. We develop a general theoretical framework and derive a practical algorithm, Distributionally Robust Relaxed Regret Optimization (DR3O). Starting from the well-known twin critics, we compare three rules to treat the critics: SAC uses its actionwise minimum and TD3 a fixed critic, whereas \drthreeo\ selects the critic under which the current actor has the larger soft hindsight gap. By only indicating "which $Q$ to give to the actor", it improves matched SAC and TD3 baselines by 5.1\%-31.7\% on average over three MuJoCo-v5 tasks. A feature-ellipsoidal \drthreeo\ variant achieves a 5.5\%-59.5\% improvement.
Chat is not available.
Successful Page Load