Learning in Context, Guided by Choice: A Reward-Free Paradigm for Reinforcement Learning with Transformers
Abstract
In-context reinforcement learning (ICRL) pretrains transformer models that adapt to unseen tasks from context alone, yet every existing ICRL method presumes scalar reward signals at both pretraining and deployment, which is restrictive in settings where rewards are ambiguous or costly to obtain. This paper asks a more fundamental question: \emph{does in-context decision-making require rewards at all, or is ordinal information, knowing which of two behaviors is better, already enough?} We show that ordinal information suffices: a transformer pretrained on comparisons alone solves tasks it has never seen, adapting from a handful of preference observations in context and without a single parameter update. We call this paradigm \emph{In-Context Preference-based Reinforcement Learning} (ICPRL) and study it under two feedback regimes, step preferences that compare two actions at the same state and trajectory preferences that compare whole trajectories. We first show that supervised pretraining remains effective with preference-only context, then introduce one preference-native algorithm per regime, requiring neither rewards nor optimal-action labels and supported by identification-level theory. Across dueling bandits, navigation, and 7-DoF robotic manipulation, ICPRL performs on par with, and at times better than, ICRL methods that consume full reward supervision, and the same holds when every preference label comes from an off-the-shelf LLM annotator, so that no reward function is involved at any stage.