RoboRef: A Foundation Reward Model Built for Robot RL, Not Offline Metrics
Abstract
Defining a reward is a long-standing problem in robot reinforcement learning (RL). The recent surge of vision-language models (VLMs), with impressive performance across domains, raises a key question: can pre-trained VLMs provide meaningful rewards and resolve the reward-definition bottleneck in robot RL? Recent work fine-tunes VLMs on large robotic datasets to investigate this, with Robometer demonstrating the most competitive performance on offline benchmarks. Yet how these reward predictions translate to downstream RL remains largely unexplored. We introduce RoboRef, which addresses two gaps shared by most VLM reward models. First, the annotated robot data these models train on mostly contain successes, which leaves these models poorly equipped to judge the failures an optimizing policy produces. We therefore re-curate RBM-1M, a large-scale robot reward dataset, to add densely annotated failure trajectories. Second, most models penalize over- and under-prediction of the reward equally, although only false positives can lead to reward hacking. We therefore introduce an asymmetric loss that penalizes reward over-estimation more harshly. We evaluate on offline metrics and downstream RL across three simulated benchmarks, with every reward model frozen and queried zero-shot, and find that scoring well on offline data does not necessarily predict downstream RL performance. RoboRef trains policies on tasks where the off-the-shelf baselines never learn at all and delivers the strongest downstream performance on average, making it, to our knowledge, the state-of-the-art VLM reward model for simulated robot RL.