NavOCR: A Dataset Generator for Navigation-Relevant Text Detection in Mobile Robots
Abstract
Textual cues provide important navigation-relevant semantic information like room names, store signs, and floor labels. However, conventional optical character recognition (OCR) systems and datasets are designed to detect all visible text, including advertisements and price tags, regardless of their relevance to navigation tasks. This characteristic of OCR can produce noisy semantic cues, making it difficult to identify which text is useful for robot navigation. A straightforward solution is to curate existing OCR datasets by retaining only navigation-relevant text and removing irrelevant annotations, but this process requires substantial manual effort. To address this problem, we introduce NavOCR, a dataset generator for navigation-relevant text detection that does not require manual curation. By leveraging conventional OCR, open map resources, and filtering with a vision-language embedding model, NavOCR identifies which text instances serve as place identifiers for robot navigation and constructs datasets from images available on the web. NavOCR aims to complete the entire process within a single day, from data collection to model training, enabling rapid adaptation to diverse environments and languages. Compact models trained on datasets generated by NavOCR achieve higher detection accuracy and substantially better computational efficiency than vision language models and open vocabulary detectors. We further show that detecting navigation-relevant text improves the quality of semantic cues used for localization and mapping in robot navigation applications.