VLAN: Vision-Language Accessible Navigation
Abstract
Accessible mobility is user-conditioned rather than scene-intrinsic, as walkable paths for one embodiment are inaccessible to another. Yet, existing navigation benchmarks evaluate traversability under a fixed and limited embodiment. We present a dataset called Vision-Language Accessible Navigation (VLAN). The dataset contains 1000 scenes, 20K navigation trajectories, and 160K keyframes, spanning diverse lighting, environmental conditions, scene semantics, and path geometries. For each trajectory, VLAN provides embodiment-specific actions and accessibility features for legged and wheeled agents, including humanoid and quadruped robots, as well as people with visual or mobility impairments. Furthermore, everyday obstacles and accessibility labels are included and categorized as Passable, Bypass, HighRisk, and Blocked. Unlike goal-only evaluations, our benchmark measures whether model outputs are both physically safe and aligned with user-specific constraints through goal-reaching, collision-aware, and accessibility-aligned metrics. We benchmark modern vision-language models and find that current systems often fail to consistently adapt obstacle relevance, path selection, and safety criteria across user profiles. These results suggest that accessibility-aware navigation requires dedicated datasets, tasks, and metrics beyond conventional embodied navigation evaluation. All data and code are open-source at https://huggingface.co/datasets/anonymousxxd/VLAN.