Reward-Estimated Hypergradient for Bilevel Reinforcement Learning with Black-Box Follower
Abstract
Optimizing the leader's policy via hypergradients (HG) in bilevel reinforcement learning (RL) typically assumes a white-box follower. To enable real-world applications, we address the black-box follower setting, where the leader must optimize solely from observed trajectories without access to the follower's true reward function. We propose Reward-Estimated Hypergradients (RE-HG), which leverages Inverse RL (IRL) to recover this unobserved reward. Because naive IRL introduces severe reward-shaping biases and numerical instabilities, RE-HG introduces a theoretical bias-cancellation mechanism that aggregates information across diverse environmental dynamics. Together with eigenvalue truncation, this acts as implicit regularization to suppress HG norm explosions. Empirical evaluations demonstrate that RE-HG estimates gradients aligned with a white-box Oracle, achieving comparable or superior performance. Notably, in sharp reward landscapes where the Oracle becomes trapped in suboptimal local minima, RE-HG's implicit regularization extracts stable global gradients, demonstrating robustness.