Benchmarking and Advancing Physical Reasoning in Large Language Models Through Agents
Abstract
Robotics is often described as the next frontier for artificial intelligence, and that integration depends on a capability most language systems were never designed for: reasoning about how objects behave in the physical world. Today that reason- ing is handled by a patchwork of narrow methods, from reinforcement learning and physics-based simulation to embodied vision-language-action models such as PaLM-E and RT-2 [Driess et al., 2023, Brohan et al., 2023], each of which works well in the setting it was built for, generalizes poorly outside it, and depends on expensive pretraining and hardware-specific pipelines. Given how quickly general-purpose large language models (LLMs) have improved at reasoning, cod- ing, and instruction following, I ask a more direct question: can an off-the-shelf LLM, wrapped in the right scaffolding, serve as a general-purpose physical rea- soning engine for robotics without any of that specialized training? To find out, I built a 200-question physical reasoning benchmark spanning ten subcategories relevant to robotic manipulation and locomotion, then used it to evaluate nine agentic architectures, from vanilla prompting and chain-of-thought reasoning to retrieval-augmented generation and tool-augmented reasoning built around exe- cutable code. The strongest of these were folded into a single hybrid agent that interleaves retrieval, staged derivation, and Python-based computation, reaching a composite score of 0.676: more than four times the vanilla baseline of 0.157 and ahead of every individually tested technique. The results point to what matters more than raw model scale: a reasoning trace that is actually visible in the output, retrieval grounded in a domain-specific knowledge base rather than parametric memory, and a calculator that never slips on arithmetic. They also show where the approach still falls short. Fluid dynamics, elastic deformation, and problems that chain several physical principles together remain stubbornly difficult, even for the best agent tested here.