Towards Optimism-Pessimism Trade-off in Model-based Offline-to-Online Reinforcement Learning
Abstract
Model-based offline-to-online reinforcement learning (RL) enables sample-efficient adaptation by leveraging offline pre-training for online fine-tuning. However, the distribution shifts between offline and online stages often hinder fine-tuning performance. Many existing methods approach this problem by adjusting the optimism-pessimism trade-off via a single-objective formulation, requiring costly online bi-level optimization. We identify this trade-off during offline training as a key challenge: optimistic policies generalize better to novel tasks by exploring out-of-distribution states and actions, while pessimistic policies remain constrained to the offline data distribution and excel on similar tasks. To address this challenge, we propose a bi-objective formulation that captures this trade-off, yielding a Pareto policy pool during offline training. These policies enable flexible selection for various online tasks. To generate the pool, we introduce Multi-Objective Soft Actor-critIC (MOSAIC), which solves bi-objective problems and constructs diverse Pareto policies. After offline training, a contextual bandit algorithm hierarchically selects the most suitable policy for fine-tuning at each online interaction step. Empirically, our pipeline, Hierarchical ParetoPolicy Pool (HiP3), achieves state-of-the-art performance on a suite of continuous-control offline-to-online benchmarks, particularly under shifted online tasks. Comprehensive ablations further clarify the roles of the policy pool and the online selection mechanism, as well as the robustness of the overall framework.