Between Policy and Controller: Post-Training Position Correction for a Frozen VLA
Abstract
A vision-language-action (VLA) policy proposes robot actions. A low-level controller turns them into robot motion. We ask whether this interface can be improved after training without changing the policy itself. We freeze SmolVLA and insert a position corrector before the controller. SmolVLA outputs seven values: requested changes in end-effector position and orientation, plus a gripper command. The corrector changes only the three position values. It uses a calibrated one-step model to compare the motion expected from the previous command with the motion that occurred, then makes a smooth, bounded adjustment to the next command. We evaluate the corrector on three simulated bowl-placement tasks from LIBERO Spatial. Both conditions use the same frozen policy, low-gain controller, and 300 matched episodes. Success rises from 195/300 (65.0%) without correction to 215/300 (71.7%) with correction. The paired gain is 6.67 percentage points (95% bootstrap interval, 4.33 to 9.33), and all ten matched groups favor the corrector. In this setting, adapting the action-controller interface improves task completion without updating the VLA weights.