One Block to Segment Them All: Unified Spatial-Channel Attention for 3D Medical Images
Abstract
Volumetric medical image segmentation requires both long-range anatomical context and precise local boundaries under tight memory constraints. Existing hybrid blocks often compute global context, local detail, and channel recalibration through separate projections and combine them only after those representations have been formed. We present CDGC-Net, whose CDGC block couples two operations. Cooperating Dual-scale Self-Attention (CDSA) partitions the projected channels between a low-rank global path and a windowed local path; for fixed projection rank and window size, its attention cost is linear in the number of voxels. Grouping Hierarchical Channel Attention (GHCA) models interactions first within channel groups and then across group descriptors. The channel query is conditioned on the spatial output, and both operations reuse the same Key projection, reducing redundant projection and avoiding a separate concatenation-based fusion layer. Experiments on Synapse, ACDC, BraTS, and LA report improvements over the UNETR++ baseline in Dice and HD95 with fewer parameters and FLOPs. Component ablations on ACDC associate the largest performance decrease with removing GHCA.