Cheap Calls Are Not Cheap Control: Success-Amortized Wall Time for On-Device Embodied Agents
Abstract
Small vision-language models are attractive for on-device agents because each1 inference call is fast. Closed-loop agents, however, choose how many calls they2 consume through their task competence and stopping behavior. We study this3 effect in an eleven-task simulated quadruped suite. In a matched seed-0 findx4 trial, a local Q4 qwen2.5vl:3b call is 2.60× faster than a cloud-model call (1.915 versus 4.97 seconds), yet the local episode is slower and fails after all 80 calls,6 while the cloud controller succeeds in 29. Across the seed-0 fixed suite, amortized7 wall-clock per successful task is 1767 seconds locally and 775 seconds for the8 best cloud reference. A local-only sensitivity over three environment seeds yields9 four successes in 33 trials and a pooled value of 1182 seconds; findx succeeds10 in one of three seeds. The checkpoint’s runtime footprint also differs from its11 file size: Ollama reports 11.0 GB allocated for a 3.2 GB model at a 4096-token12 context. This narrow case study—one quantized model, one device, and unseeded13 decoding—is a concrete seed-0 counterexample, not a model-family ranking.