RLVR Makes Agents Win More, but Not Necessarily Reason Better
Abstract
Reinforcement learning from verifiable rewards (RLVR) can substantially improve the task performance of language and vision-language agents. Yet outcome gains do not reveal the quality of an agent's visible reasoning. In interactive settings, this concern extends across turns as observations and action outcomes accumulate, raising an urgent need to assess how reliably an agent maintains a running account of the ongoing interaction, especially when its traces are used for human understanding or auditor oversight. To address this need, we systematically evaluate reasoning traces from interactive agents throughout RLVR, with \emph{reasoning trace reliability} as a multidimensional target operationalized by three rubrics: grounding in the current observation, coherence between reasoning and the selected action, and temporal consistency with interaction history. Empirical results show that task performance and trace reliability may improve together across training, yet within a fixed checkpoint, traces from higher-performing trajectories are often not materially more reliable. Reliability also tends to decline toward later interaction turns, even in trajectories that ultimately succeed. Qualitative trace audits further show agents often carrying stale or unsupported task-state beliefs into later-turn reasoning after failed actions. Together, the analyses clarify RLVR’s limits in cultivating reliable reasoning beyond surface plausibility throughout an interaction and identify where multi-turn oversight and efforts to recover trace reliability should concentrate.