SimVLA: Attributing Gains in VLA Models Through Controlled Ablation
Yuankai Luo ⋅ Woping Chen ⋅ Tong Liang ⋅ Zhenguo Li
Abstract
Vision-Language-Action (VLA) models often bundle architectural changes with different pretraining data, backbone scales, and optimization recipes, obscuring what actually drives progress. We introduce SimVLA, a deliberately minimal VLA—a standard vision-language backbone with a lightweight continuous-action head—as a controlled testbed for this attribution problem. Across LIBERO, CALVIN, WidowX, and Google Robot, one-knob-at-a-time ablations show that training dynamics are the largest measured driver, producing performance swings of $\Delta$$\approx$54-89\%, while task configuration and architecture have smaller effects under matched settings. Despite using a small backbone and no additional large-scale robot-data VLA pretraining before benchmark-specific training, SimVLA reaches 98.6\% average success on LIBERO, remains competitive across the other benchmarks, and transfers to held-out real-robot scenes. These results suggest that simple, carefully controlled VLA baselines remain underexplored and provide a calibration point for future architectural claims within the VLM-encoder plus lightweight-head family.
Chat is not available.
Successful Page Load