When Delay Meets Distribution Shift: Evaluating Asynchronous Vision-Language-Action Policies
Abstract
Asynchronous execution allows continuous movement in a robot while its policy computes the next action chunk in the background. However, this form of execution comes with two risks: the actions remaining in the queue could be predicted from a stale observation, and the queue may be exhausted before the inference finishes. This paper examines whether distribution shifts make inference delay more harmful. We use a fixed LIBERO-fine-tuned π0.5 checkpoint and compare matched LIBERO and LIBERO-Plus conditions using a discrete-event executor. For each request, it combines measured inference latency with an added logical delay. We select +200 ms using in-distribution (ID) trials only, then apply it to both ID and out-of-distribution (OOD) conditions. Across the pooled Naive Async and Real-Time Chunking (RTC) screen, we do not observe a larger delay-related success drop under OOD than under ID. Action coverage shows a clearer difference. In ID-only RTC tests at +200 ms, 10 configured actions give 6/15 successful episodes, compared with 14/15 at each of 20, 25, and 30 actions. Our eight-seed held-out follow-up identifies one task–variant pair with a larger estimated OOD delay penalty. Its uncertainty intervals include zero, however, and results from other tasks and layouts are mixed. We therefore suggest checking action coverage against the tested delay before attributing failures to distribution shift. Our complete implementation and benchmark suite are publicly available at: https://anonymous.4open.science/r/vla-latency-benchmark-23CB/README.md.