SynerVLA: Exploiting Embodied Execution Phases for On-Device Dual-System VLA Acceleration
Abstract
Dual-system visual-language-action models, which integrate high-level planning (System 2) with instant control (System 1), are promising for embodied AI, but their deployment is hindered by the high computational cost of processing continuous visual streams. Existing acceleration methods are one-sided, focusing only on optimizing the VLM (System 2), which not only limits speed gains but also risks compromising the deep reasoning capabilities it is meant to provide. In this paper, we introduce SynerVLA, a plug-and-play framework for accelerating dual-system VLA models. Its core idea is to leverage the distinct phase information in embodied execution, rapid approaching and fine-grained manipulation, to exploit both spatial and temporal redundancy in visual streams throughout the entire execution pipeline. SynerVLA accelerates dual-system VLA inference by first selecting key visual tokens via text-action fusion, then dynamically adjusting the token pruning-reuse ratio through phase-aware feedback, and finally accelerating both System-2 VLM and System-1 diffusion Transformer with a dual cache reuse mechanism. Evaluations on representative platforms demonstrate that SynerVLA delivers up to a 1.93× speedup and a 26\% higher control frequency, with only a negligible impact on task success rate.