Robust Offline Reinforcement Learning against Out-of-Distribution Dynamics in Autonomous Driving
Abstract
Real-world autonomous driving data inherently exhibits out-of-distribution (OOD) dynamics, as diverse interactions and unpredictable human behaviors cause vehicle trajectory distributions to vary significantly across different scenarios, even for similar road structures. Such OOD dynamics typically induce uncontrollable extrapolation errors in neural networks. We observe that in offline reinforcement learning (RL), these errors manifest as heavy-tailed estimates of action values with significant biases, which existing robust methods fail to address, hindering effective generalization. To address this problem, we propose a robusT offline RL algorithm under heavy-tailed action-valUe eSTimates (TRUST), which models perturbed states associated with heavy-tailed action-value estimates to effectively mitigate such estimation biases. Specifically, TRUST uses a first-order gradient-regularized target to model perturbed states that approximate data under worst-case dynamics. It then introduces a statistical measure to identify perturbations that induce heavy-tailed action-value estimates. By re-weighting the training objective to emphasize correcting these biased estimates, TRUST can reduce the heavy-tailed behavior in action values for robustness against OOD driving data. We further derive an error bound that justifies the gradient-regularized target as an approximation to the robust Bellman target. Extensive experiments on three large-scale real-world datasets show that TRUST achieves strong overall performance across all benchmarks, with a general performance improvement of 41.21\%.