Retrieve, Don’t Retrain: Extending Vision-Language-Action Models to New Tasks at Test Time
Abstract
Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute. We show that this target-side per-task adaptation cost can be replaced by retrieval. Our retrieval-augmented policy is trained once on paired demonstrations from the target embodiment (query) and a cheaper embodiment (pool, e.g., human-hand video), then frozen. New tasks are added at deployment by appending pool-side demonstrations to a retrieval pool, and the frozen policy conditions on retrieved trajectories at every control step, so new tasks are absorbed by indexing data rather than updating parameters. Fine-tuning is then needed only to take on a new, unseen embodiment, not for each new task. Retrieval helps on both an action-only VLA and Cosmos Policy, a video-generation-based world-action model (WAM), but its effect is markedly stronger on the WAM: retrieval supplies coarse task progression, while the WAM's future-image objective provides a visual consistency signal that, on PushT, improves unseen-task success only when retrieval is present. On a cross-embodiment PushT variant, retrieval raises unseen-angle success from 6.0% to 34.9% (n = 350 rollouts). On RoboTwin 2.0, the frozen policy reaches 31.5% on five held-out tasks, statistically indistinguishable from a policy fine-tuned on each of them (24.0%, p = 0.09) and the only method that is strong on seen and unseen tasks at once (37.5% vs. 25.8% overall for the strongest baseline, p < 0.001). On a real robot, adding ten human-hand demonstrations per held-out task raises success from 10% to 80% on placing a bottle without any robot data for that task.