Where Edge VLA Latency Actually Goes: Measured Scaling Laws and Staleness Admissibility for a 450M Vision–Language–Action Model
Prarabdh Misra ⋅ Mohd Ariful Haque
Abstract
Most work on making vision–language–action (VLA) models fast enough for edge robots targets the vision encoder, because pre-fill is reported to take over 90% of forward-pass time. We measured a 450M-parameter flow-matching VLA and found the opposite. The action expert takes 67% of the single-stream budget, and it is the expert that batches well; pre-fill barely does. Three findings follow. First, latency tracks encoder pixel work (images $\times$ resolution), not how many visual tokens the language model sees: a two-parameter fit, $\mathrm{pre\text{-}fill} \approx 23 \mathrm{ms} + 464 \mathrm{ms}/\mathrm{megapixel}$, predicts a held-out setting to 0.1%, and batching headroom follows the same quantity, so three cameras at 256 px batch better than one at 512 px. Second, caching the vision prefix across frames – the standard dual-rate trick – is safe only while the scene stays still. Under an action-error ceiling matched to INT8 quantization, no reuse depth survives any of the 24 moving conditions we tested, on either of two stimuli spanning the range of spatial coherence. Third, and not by design, the coherence of that test stimulus decides which mechanism the experiment reports. On i.i.d. noise – the default a latency harness reaches for – the penalty is flat in reuse depth, which reads as stale vision costing a fixed price while stale proprioceptive state compounds, and yields a crisp rule. On a $1/f$ field both terms compound and the gap between the two reuse variants falls from $3.24\times$ to $1.40\times$: the mechanism was the stimulus running out of information. The headline survives, the explanation does not, and we report both grids. Finally, we report action error as a fraction of the action standard deviation rather than cosine similarity, because the two disagree: INT4 quantization scores 0.977 cosine similarity while costing $6.9\times$ the action error of INT8, a gap that aggregate metrics hide.
Chat is not available.
Successful Page Load