Deferral Is Not Free: The Cost of Small-to-Large Handoff in Embodied Agents
Abstract
Cascaded agent stacks run a small model locally and escalate hard decisions to a stronger one. Calibrated cascades price that escalation in tokens or dollars, resting on the premise that deferring is safe because only accepting commits an error. In an embodied setting the price is also paid in time: the world keeps moving during the round trip, and the agent is not acting while it waits. We charge deferral against a world clock in a dynamic gridworld where fire consumes objects on per-object deadlines. We find a critical round trip Delta* beyond which escalating leaves the agent worse off than never escalating, measured at 4, 10, and 12 ticks across three hazard velocities. Decomposing the cost, lost acting time exceeds plan staleness at every measured condition, by a median factor of 23, though both channels are significant at 500 seeds. Escalation granularity matters more than latency: at zero delay, action-level escalation recovers 8-21% of the available quality gap while committing to a returned plan recovers 89-105%. Across five on-device models from 0.5B to 3B, valid-plan rate climbs with scale to 93% while none shows a measurable quality advantage over a greedy heuristic. Escalation must clear two bars, a real quality gap and a round trip under Delta*, and the whole family clears the second while none clears the first.