Refinement versus Refresh: Compute Allocation for Closed-Loop Robot Policies
Abstract
Robot policies can spend substantial inference compute refining actions before execution, but it is unclear whether this extra computation improves closed-loop control. We study this deployment trade-off in one frozen SmolVLA policy, combining a paired, 12-configuration discovery study with a prespecified comparison on eight LIBERO tasks excluded from configuration selection. At a fixed two-action prefix, reducing denoising from five steps to two increases success from 122/160 to 139/160: a paired difference of 10.6 percentage points with a task-conditional, descriptive 95% interval of 4.4–16.9 points. The observed direction agrees with the frozen prediction of 4.0 points, but its magnitude is larger. Relative to the same comparator, reference-calibrated dense arithmetic is 5.9% lower per control step; mean arithmetic per attempt is 41.1 versus 48.7 TFLOPs, including failures. The 480-attempt discovery study varies both refinement and execution prefix, and full-generation calibration assigns 82.1% of native counted arithmetic to denoising-independent work. A simple cost-aware rule matches the fitted response model’s configuration choices at three budget ceilings. For this checkpoint, the separate-task comparison shows that reducing refinement can improve closed-loop success while lowering computation. These findings motivate a broader question: when should inference compute be spent refining an action prediction, and when should it instead be spent obtaining a new observation-conditioned prediction?