Rethink Action Chunking in VLA Through Human Motor Control
Abstract
Action chunking is widely used in recent vision–language–action (VLA) systems to mitigate inference latency and improve temporal consistency. Yet, open-loop execution can reduce reactivity and precision, and cross-chunk coherence remains difficult to guarantee. We revisit action chunking through its neurobiological origins and organize current shortcomings around three axes: VLA architecture, representations of sequential actions, and mechanisms for sampling and executing chunks. A case study further illustrates the coherence–reactivity trade-off across representative sampling and execution choices. Building on these observations, we outline brain-inspired directions to improve the reactivity, coherence, and safety of chunking-based VLAs in dynamic and complex tasks beyond quasi-static tabletop manipulation. This position paper aims to broaden how the community thinks about action chunking and to motivate redesigns better aligned with future physical intelligence.