OLA-Place: Cross-Modal Place Recognition without Global Descriptors
Abstract
Cross-modal place recognition aims to retrieve a target 3D location from a spatial map using a natural language description. Most existing methods follow a global-descriptor matching paradigm, in which the text query and each 3D scene cell are independently compressed into single vectors and compared by global similarity. Although simple and widely adopted, this paradigm tends to discard the fine-grained object correspondences that are essential for language-guided localization. In this paper, we challenge the necessity of global descriptors and propose OLA-Place, an Object-Level Alignment framework for cross-modal place recognition without global descriptors. Instead of representing a query or a scene cell as a single holistic embedding, OLA-Place formulates place recognition as set-to-set semantic alignment between textual object mentions and 3D object instances. The framework contains three key modules: ObjectSet Encoder (OSE), which extracts object-level representations from both language descriptions and 3D scene cells; Object Message Encoder (OME), which injects intra-set object-context information while preserving object-level granularity; and Masked Max Alignment (MMA), which computes the query-cell matching score by aligning each textual object mention with its most similar valid 3D object. This formulation is permutation-invariant, naturally handles variable-size object sets, and requires no object-level correspondence annotations. Extensive experiments show that object-level alignment alone significantly outperforms global-descriptor-based methods, demonstrating that cross-modal place recognition is better understood as an object-level semantic correspondence problem rather than a global-embedding retrieval problem. Our code is available at: https://github.com/Anonymous09871745/OLA-Place.