Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA
Abstract
Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay intervention when unexpected scene changes require renewed high-level reasoning. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot, computed from the specialist's existing forwards under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive, enabling TUD to invoke the generalist only when the cached plan becomes unreliable, and separates success from failure more reliably than prior uncertainty signals. The signal needs no manually labelled phase boundaries or auxiliary uncertainty model, and is computed from forwards the architecture already runs. On VLA-Arena, TUD finds a more favorable cost--success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.