RelaVPR: Relation-Based Knowledge Distillation for Efficient Visual Place Recognition
Abstract
Visual Place Recognition (VPR) aims to identify and retrieve geographic locations from visual observations. With the advent of Visual Foundation Models (VFMs), recent VPR methods have primarily adopted them as backbones to exploit their strong generalization capabilities, effectively mitigating challenges such as perceptual aliasing and long-term appearance variations. However, their large parameter sizes and high inference latency impose substantial hardware demands, significantly hindering practical deployment. To retain the generalization strength of VFMs while enabling efficient VPR, we revisit this problem from the perspective of Knowledge Distillation (KD). Specifically, we conduct a systematic investigation of various KD losses for VPR from both theoretical and empirical standpoints. Our analysis reveals that relation-based KD methods consistently achieve superior performance and efficiency, which we attribute to their larger feasible solution space and better alignment with the VPR objective. Building upon this insight, we propose a relation-based KD framework, termed RelaVPR, which is further enhanced along three key dimensions: 1) dynamically filtering noisy knowledge, 2) improving hard-sample learning, and 3) mitigating gradient interference in multi-teacher distillation. Compared with state-of-the-art (SOTA) KD-based VPR methods, RelaVPR requires no additional fine-tuning while delivering higher performance, faster inference, and a more compact model. Moreover, RelaVPR establishes a stronger performance--efficiency frontier than existing SOTA VPR approaches across multiple popular benchmarks. Codes and weights will be released.