Think During Training, Act Without the LLM: Distilling LLM Guidance for Long-Horizon Robot Navigation
Abstract
Foundation models provide powerful reasoning priors for robot decision making, but repeatedly querying them during policy execution can introduce substantial latency and computational overhead. We investigate an alternative paradigm: use foundation-model reasoning as guidance during reinforcement learning, while learning a specialized policy that operates independently of the model at test time. We study this problem in long-horizon legged navigation under partial observability. Our framework combines a three-level hierarchical reinforcement learning (HRL) framework. A low-level policy tracks body-frame velocity commands, a mid-level policy drives the robot towards local subgoals, and an event-driven high-level policy selects grounded candidate subgoals generated from an online occupancy memory. Our core idea is to use a small local LLM as a training-time guidance and distill only its beneficial exploration decisions. The LLM and RL policy select from the same belief-map-based candidate representation. Each LLM proposal is first grounded by a geometric verifier, while advantage-filtered distillation retains only actions whose outcomes exceed the current policy's expected value. Because these outcomes reflect the complete hierarchy, the learned high-level policy favors subgoals compatible with the executing controller. Trained only on a simple U-shaped maze, the LLM-guided HRL policy generalizes without fine-tuning to three unseen maze layouts, achieves the strongest overall learned performance and the highest average SPL compared to baseline. At test time, decisions from LLM-guided HRL require approximately 1 ms, over three orders of magnitude faster than direct LLM-in-the-loop planning. These simulation results suggest that foundation-model reasoning can serve as effective training-time guidance for learning efficient specialist policies that no longer require foundation-model inference during execution.