SpanFormer: Multi-Level Adaptive Sparsity for Object Detection in High-Resolution Wide Shots
Abstract
Object detection in high-resolution wide (HRW) shots, where a single image can span gigapixel resolutions and kilometre-scale fields of view, suffers from extreme foreground sparsity, dramatically varying foreground density from sparse to crowded scenes, and object clusters whose spatial extent ranges from dozens to thousands of pixels within the same dataset. State-of-the-art sparse vision transformers for gigapixel detection select a fixed top-k fraction of windows by hand-crafted variance scores, which over-promote textured but object-free background such as foliage and patterned facades, cannot adapt the keep ratio to the wide range of crowd density, and confine each kept window to a 7x7 receptive field that is far smaller than the spatial extent of many real-world object clusters. We propose SpanFormer, a sparse vision transformer that broadens the computational span at three complementary granularities: (i) Prototype Routing replaces variance with learnable foreground/background prototypes regularised by a balance constraint that prevents prototype collapse, widening the semantic span of the scoring criterion; (ii) Dynamic Top-k predicts a per-image keep ratio with a lightweight selector, making the compute span elastic so that crowded scenes receive proportionally more compute while near-empty scenes are skipped; (iii) Cluster Memory Attention performs union-find clustering over the kept windows by joint cosine similarity and Chebyshev spatial radius and augments each window's K and V with cross-window memory drawn from its cluster, extending the attention span beyond the local window with all projection parameters shared with the local branch and therefore no added projection cost. On the PANDA gigapixel benchmark, SpanFormer improves AP50 from 78.0% to 80.3% over the SparseFormer baseline while reducing backbone FLOPs by over 75%.