Rethinking Infrared Small Target Detection: A Foundation Driven Efficient Paradigm
Abstract
While large-scale visual foundation models (VFMs) exhibit strong generalization across diverse visual domains, their potential for infrared small target (SIRST) detection remains largely unexplored. To fill this gap, we systematically introduce frozen VFM representations into the SIRST task and propose a Foundation-Driven Efficient Paradigm (FDEP), a general framework compatible with diverse VFMs and SIRST networks, which improves detection accuracy without additional VFM-related inference overhead. Specifically, a Semantic Alignment and Modulated Fusion (SAMF) module is designed to achieve dynamic alignment and deep fusion of the global semantic priors from VFMs with task-specific features. Meanwhile, to avoid the inference-time overhead introduced by VFMs, we propose a Collaborative Optimization-based Implicit Self-Distillation (CO-ISD) strategy, which enables implicit semantic transfer between the main and lightweight branches through parameter sharing and synchronized backpropagation. In addition, to unify the fragmented evaluation system, we construct a Holistic SIRST Evaluation (HSE) metric that performs multi-threshold integral evaluation at both pixel-level confidence and target-level robustness, providing a stable and comprehensive basis for fair model comparison. Extensive experiments demonstrate that the SIRST detection networks equipped with our FDEP framework achieve state-of-the-art (SOTA) performance on multiple public datasets. Our code will be open source.