When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning
Elias Hossain ⋅ Mohammad J Basher ⋅ Ivan Garibay ⋅ Ozlem Garibay ⋅ Niloofar Yousefi
Abstract
Offline RL can learn effective policies from fixed datasets, but deployment objectives often shift after training, while the trained actor may be frozen due to data, cost, or governance constraints. We study deployment-time adaptation of frozen offline actors using Product-of-Experts (PoE) composition with a goal-conditioned prior. Our main practical finding is graceful degradation rather than universal improvement: with degraded or random priors, precision-weighted PoE remains anchored to the frozen actor, whereas additive and prior-only adaptation often collapse, and a KL-budget selector typically reaches a near-oracle operating point. We also clarify that, for diagonal-Gaussian actors and priors, PoE with coefficient $\alpha$ induces the same deterministic policy as KL-regularized adaptation with $\beta=\alpha/(1-\alpha)$, with posterior covariance differences cancelling under deterministic action selection. Across D4RL MuJoCo and AntMaze diagnostics, results reveal an actor-competence ceiling: composition helps or preserves performance when the frozen actor is sufficiently competent, but cannot recover when the actor itself fails. Swept CQL/IQL-guided and SF/GPI baselines are competitive in some settings, but do not eliminate this regime boundary. Thus, PoE/KL-Reg is best viewed as an actor-anchored safety layer for deployment-time objective shift, not as a general return-maximization method, with effectiveness governed by actor competence rather than knob choice.
Chat is not available.
Successful Page Load