DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
Abstract
Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally. However, many states provide limited discrimination signal because plausible actions are similar, while a sparse set of decision-critical states admit meaningfully different actions that can substantially affect downstream outcomes. We therefore propose DiVeR, a decision-criticality-weighted verifier that estimates state importance from the dispersion of sampled action representations. Using this signal, DiVeR places greater emphasis on decision-critical states during training, focusing verifier learning where action selection matters most. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves verifier-guided action selection with negligible inference overhead.