WebSpatial: A Benchmarking Framework of Spatial Intelligence via Web-based Runtime Environments
Abstract
Spatial intelligence is a fundamental capability of embodied Artificial Intelligence systems, and its reliable assessment and optimization require well-designed benchmark. Existing benchmarks for spatial intelligence largely rely on video datasets or procedurally generated data, which limits data scalability and ecological validity. To address these limitations, we propose to automatically extract spatial information directly from executable environments. Specifically, we design WebSpatial, a benchmarking framework built upon web runtime environments, which enables automatic acquisition of spatial data by parsing underlying code and interaction processes. It performs runtime scene graph parsing on web 3D applications to automatically extract visual attributes, world coordinates, and user interaction event sequences of objects, while constructing a dedicated spatial computation tool library. The modular generation pipeline integrates LLM-based task parsing and planning and tool-assisted execution to automatically generate question-answer pairs. We conduct experiments in method effectiveness validation and benchmark evaluation. For validation, we demonstrate that WebSpatial can reliably construct diverse and scalable spatial tasks grounded in both static and dynamic scenes, covering spatial perception including color, shape and quantity, and spatial reasoning including relation assessment, direction transformation and mental rotation. We develop a small-scale benchmark named WebSpatialBench based on WebSpatial, and use it to evaluate 10 multimodal large language models, revealing their strengths and limitations in spatial problem-solving.