Omni-SpikeDet: A Spiking Open-World Detector with Dynamic Text–Image Alignment
Abstract
Open-vocabulary object detection requires recognizing both known and novel categories beyond a fixed label space. Existing open-vocabulary detectors have made strong progress by integrating visual features with language embeddings, but they are predominantly built on dense ANN-based architectures. In contrast, spiking neural networks (SNNs) have been studied for sparse visual computation, yet their use in open-vocabulary detection remains largely unexplored. This paper presents Omni-SpikeDet, a spiking open-vocabulary detector that investigates how text-guided visual recognition can be realized in a spiking framework. Omni-SpikeDet combines spike-based visual feature extraction with text-guided multi-scale alignment. It uses a spiking convolutional module with integer-spike training and spike-based inference to reduce the mismatch between continuous vision-language supervision and discrete spiking representations. It further introduces a Cross-Scale Text-Image Interaction (CTI) module that modulates multi-level spiking features with language embeddings for prompt-guided detection. Experiments on COCO and LVIS show that Omni-SpikeDet achieves competitive open-vocabulary detection performance with a compact model, obtaining 47.2 AP on COCO and 27.9 AP on LVIS with 4.2M parameters. Under a standard theoretical SNN energy model, Omni-SpikeDet has an estimated energy cost of 19.8 mJ, suggesting a favorable accuracy-efficiency trade-off.