The Optimal Brain Size for a Body
Abstract
A larger neural network can predict better yet control a robot worse when inference takes longer. This trade-off matters for embodied AI, including vision-language models and vision-language-action policies. We quantify the cost of inference latency. For optimal linear-quadratic control under white actuator noise, an exact identity shows why the leading delay cost is linear: fresh disturbances arrive during computation, and even perfect prediction retains this leading cost. With a single actuator, relative cost is approximately feedback bandwidth times information age when crossover lies above the plant dynamics. The same 0.85 ms delay raised cost by 0.7, 4.4 and 29% in three simulated bodies, against 0.6, 4.4 and 26% predicted. Prospectively testing 21 camera convolutional neural networks, pricing delay selected a network within 0.6% of the best in all 12 settings; choosing the lowest undelayed cost was up to 26.9% worse. With faster cameras, the simple bandwidth-based price selected the best in 12 of 16 settings and came within 1.1% in the rest. To first order, state-independent execution is priced by mean information age. Extensions examine other linear gains and nonlinear policies. Model capacity must therefore be matched to the dynamics of the body it controls.