Invariant Reward Learning Beyond Spurious Correlations
Abstract
This paper explores bias mitigation in Reinforcement Learning from Human Feedback (RLHF) by reducing shortcut dependence in reward modeling. RLHF is a training paradigm that aligns a pretrained language model with human preferences by incorporating human feedback into reinforcement learning. In the RLHF pipeline, a reward model is trained to distinguish preferred responses from rejected ones and assign higher scores to preferred responses. This reward model plays an important role by providing the optimization signal for reinforcement learning, guiding the language model toward preferred behavior. We investigate a biased setting in which the reward model suffers from prompt-derived shortcut attribute, making reward scores unrepresentative of true response quality. This bias arises when prompt-derived variables, such as sentiment or topic, influence both the generated response and the assigned reward. For example, responses to positive prompts may receive higher reward scores despite being equally correct and relevant as responses to neutral or negative prompts. We formulate this bias as a prompt-induced shortcut problem in the reward modeling stage of RLHF and study how reward scores can become overly sensitive to prompt-derived variables. To address this challenge, we propose an adversarial learning approach that discourages the reward model from relying on prompt-derived shortcut signals. We evaluate our framework using different prompt-derived variables and measure how effectively reward-score disparities across these variables are reduced across multiple RLHF datasets. The results show that our method effectively reduces the reward model’s reliance on prompt-derived shortcut attributes that cause bias, compared to baseline approaches.