Demonstration Diversity Is Not Free at a Fifty-Episode Adaptation Budget
Abstract
A practitioner adapting a pretrained vision-language-action model to a new embodiment has a fixed demonstration budget and must decide how to spend it. We pre-registered a comparison of four collection strategies at 50 demonstrations each, repetition (Clean), position diversity, recovery demonstrations and color diversity,holding the model (SmolVLA), task, hardware and training fixed: 600 rollouts across five evaluation cells and two seeds on a low-cost SO-ARM101. Both registered comparisons were null, one at a floor (0 of 120 held-out positions) and one at a ceiling. Trajectory telemetry explains both, and a third the success rates recorded but never explained: every other condition outscored the position-diverse policy by 2.4 to 4.6 times, yet that policy reached the lowest training loss of any run. Policies trained at a single position execute a fixed sweep that stops covering the cube within one inch. The position-diverse policy aims at the cube inside its training hull, yet never leaves its home pose in a fifth of episodes. A follow-up policy trained at two positions with 25 demonstrations each executes perfectly at both and, everywhere else, returns to one of its two trained bearings: density buys execution, not interpolation. Policies also inherit their demonstrator: a demonstrated mid-carry drop is reproduced within one degree of its location, and demonstration velocity orders completion time. Success rate alone would have reported no effect anywhere, which is the practical hazard at this end of the ecosystem.