From Verified Rewards to Acting Agents: Diagnosing Transition Knowledge in Physical Control
Abstract
Verified rewards tell an agent how well an action performed, but do not directly teach it how that action changed the environment. We ask whether this distinction helps explain when pretrained physical knowledge becomes useful for control. We study the same open-weight reasoning model in thermal-storage scheduling and multi-zone ventilation and cooling control, using reinforcement learning with verifiable rewards (RLVR). We test transition knowledge by asking the model to predict what would happen under alternative actions. In storage scheduling, the model could predict the tested action effects, and RLVR made it choose useful actions more often. The gains also transferred to new storage conditions and partly to battery scheduling. In ventilation and cooling control, the model struggled to predict action effects, and exact rollout rewards improved neither these predictions nor overall control performance. This contrast suggests that transition knowledge may help explain the different training outcomes, though the task comparison does not establish a causal link. We propose teaching action effects through supervised learning, then testing whether RLVR turns that knowledge into better control.