Spatial Temporal Reasoning Models for Accurate and Efficient Vision-Based Localization
Abstract
This paper explores vision-based localization through an approach that mirrors how humans link first-person and bird's-eye perspectives when navigating their world. We introduce two sequential generative models, VAE-RNN and VAE-Transformer called Spatial Temporal Reasoning Models (STRMs), that transform first-person perspective (FPP) observations into global map perspective (GMP) representations to create precise geographical coordinates. Unlike retrieval-based methods, our approach frames localization as a generative task, learning direct mappings between perspectives without dense satellite image databases. We evaluate the models on two GPS-challenged real-world environments: a university campus navigated by an unmanned ground robot and an urban downtown area navigated by a Tesla sedan. Our method surpasses three representative cross-view baselines (VIGOR, TransGeo, and S2GP) on both datasets. STRM has faster responses than these baselines and runs on a modest onboard CPU. The onboard localization with short latency and low power makes STRM attractive for on-device applications in challenging environments.