Post-Selection-Safe Pessimistic Utilities for Offline Multi-Objective Reinforcement Learning
Mingxi Hu ⋅ Meiling Yu
Abstract
We study post-selection-safe deployment in finite-horizon tabular offline multi-objective reinforcement learning with full vector rewards, nonnegative linear scalarization, and preferences revealed at deployment. A fixed dataset may already have been reused to construct a data-dependent policy library $\Pi_N$; the remaining problem is to return auditable lower-confidence utilities for future queried preferences and for small policy menus. We introduce the Pessimistic Utility Oracle (PUO), a coordinatewise pessimistic vector-value construction. On one dataset-level tabular model-confidence event, PUO lower-bounds every data-dependent policy under every $w\in\mathcal{W}\subset\mathbb{R}_+^d$, without preference discretization or an additional library-size union bound beyond the model event. PUO yields a queried-preference planner and SPS++, a greedy policy-set method. The main SPS++ theorem evaluates the executable rule that deploys the in-set policy with largest PUO score, and decomposes its loss into coverage, greedy approximation, and preference-sampling terms. Controlled tabular audits validate these statistical effects; trajectory-count and linear-feature appendices are presented only as bridges beyond the main generative tabular theorem.
Chat is not available.
Successful Page Load