Bayesian Decision-Time Inference for In-Context Reinforcement Learning from Suboptimal Data
Abstract
In-context reinforcement learning (ICRL) promises rapid adaptation without parameter updates, but standard supervised objectives often fail when pretraining data is generated by suboptimal behaviour policies. In these regimes, logged actions are unreliable labels while rewards still provide value-relevant information. To address this, we introduce SPICE, a Bayesian decision-time inference method that shifts online ICRL from action-logit prediction to approximate posterior inference over action values, requiring neither expert action labels nor algorithmic learning traces. SPICE learns a task-conditioned value prior with a transformer value ensemble and, at test time with parameters frozen, fuses this prior with kernel-weighted context evidence via a closed-form Gaussian fusion update. The resulting estimates drive a posterior-UCB controller, enabling principled online exploration and adaptation without gradient updates. A stochastic-bandit analysis shows logarithmic regret growth for the fixed-prior controller under scheduled exploration, while quantifying the additional early cost caused by inaccurate prior estimates. Across bandits, Darkroom, image-based MiniWorld, and continuous building control, SPICE adapts more effectively from suboptimal data than supervised ICRL baselines.