Offline Inverse Reinforcement Learning with Unified Diffusion Planning
Abstract
Offline Inverse Reinforcement Learning (IRL) aims to recover a reward function and imitate expert behavior solely from offline demonstrations. While recent Offline IRL approaches employ approximate dynamics to mitigate distribution shift, their performance is constrained by the unstable minimax framework as well as simplified representation and utilization. To address these issues, we propose Offline Inverse Reinforcement Learning with Diffusion Planner (OIDP), which leverages diffusion models to achieve stable Offline IRL. OIDP formulates a model-based conservative Q-optimization through inverse soft Q-learning, which is provably concave in Q-space with a controllable performance gap. Accordingly, we introduce a unified diffusion planner that models dynamics, uncertainty estimator, and policy, it captures multimodal distributions, directly tightens the objective gap bound, and at execution performs trajectory-guided policy sampling using the internalized dynamics knowledge. Experimental results on D4RL benchmarks demonstrate that OIDP outperforms state-of-the-art offline approaches.