TrajReasoner: Metric Trajectory Reasoning in VLMs for Physical-World Embodied Agents
Yue Hu ⋅ pengcheng chen ⋅ Brandon Chu ⋅ Yubin Zhou ⋅ Zhiqiang Lao ⋅ Rong Liu ⋅ Jintang Xue ⋅ Andrew Feng ⋅ Xiyun Song ⋅ Heather Yu
Abstract
Understanding the physical world through geometry is a prerequisite for embodied agents, which require metric spatial self-localization to act effectively; yet vision-language models (VLMs), now widely adopted as foundation models for such agents, remain unreliable at metric camera geometry. We present TrajReasoner, a two-stage framework that trains VLMs to estimate camera pose trajectories via geometric chain-of-thought reasoning, combining supervised fine-tuning with Group Relative Policy Optimization (GRPO) reinforcement learning on geometry-grounded rewards. Trained on 684,305 frames from RealEstate10K, our model achieves $3.2\times$ and $1.7\times$ higher structured output parse rates over a base 8B VLM and Gemini 3.1 Pro, respectively, and surpasses Gemini in translation accuracy by over $6\times$. These results demonstrate that structured training can instill metric geometric reasoning in VLMs, advancing them toward geometry-aware world models that ground physical-world understanding for embodied agents.
Chat is not available.
Successful Page Load