Geometry Is Not a Prompt: What Physical-World Agents Should Compute Instead of Infer
Abstract
Physical-world agents often combine two different kinds of reasoning: interpreting flexible language and computing quantities determined by the physical world. Treating both as prompting problems can be unreliable, particularly when correctness depends on camera pose, field of view or geometric projection. We formulate mechanism allocation: a task is divided into sequential sub-tasks, and each is assigned either to an LLM or to a specialized tool based on its expected error, latency, cost, and flexibility. We use these quantities to compare complete agent designs and study how errors in their estimation affect the selected allocation. We apply the framework to asset-aware retrieval of images in real industrial applications, where an agent aims to retrieve images containing a given asset under distinct requirements such as location in the frame. The study compares direct prompting, prompts augmented with equations, model-generated code, and deterministic geometric tools. The aim is not to replace the LLM, but to retain it for decisions requiring language understanding while computing explicitly what is best computed rather than inferred.