One Scan is Enough: Demonstration-Free Adaptation for Language-Guided Navigation
Abstract
Deploying a language-guided navigation policy in a new building usually means watching it fail. Public-benchmark VLAs degrade scene by scene as layout, appearance, and language drift, and the standard fix of teleoperating fresh trajectories on site does not scale. We replace human demonstration with a single brief iPhone scan. Scan2Instruction compiles that scan into a complete training environment, automatically synthesizing executable navigation episodes with grounded instructions. PRPT-KS then adapts the policy through a closed-form Pareto utility read directly off the reconstructed scene, with no learned critic and no human labels. Across five real environments, the two designs together drive our RealNav to an 89% real-robot success rate, almost doubling the strongest VLA baseline at 48%, and still achieve 87% in unscanned regions of the same buildings. Our work turns demonstration-free scene adaptation for language-guided navigation from an aspiration into a working reality. All code and model will be released open-source upon acceptance.