Token Mixer Placement Matters: A Systematic Encoder–Decoder Ablation Study of Mamba for Brain Tumour Segmentation on BraTS-Africa
Abstract
Accurate 3D brain tumour segmentation requires balancing fine-grained local feature extraction with long-range spatial dependency modelling. While Convolutional Neural Networks (CNNs) capture local features well, they struggle with global context. Transformer-based architectures address this via self-attention but incur prohibitive quadratic memory costs in 3D settings, limiting their utility in resource-constrained clinical environments. State Space Models (SSMs) like Mamba offer a linear-time sequence modelling alternative. However, their optimal placement within volumetric segmentation networks remains insufficiently explored, and standard benchmarks frequently fail to reflect the imaging heterogeneity of real-world practice. To address this, we present a systematic encoder-decoder ablation of Mamba-based architectures using a lightweight, unified MetaUNETR framework. We evaluate our models on BraTS-Africa 2025, a multi-centre dataset comprising 146 multiparametric MRI cases representing diverse sub-Saharan African clinical environments. After initializing with 2D ImageNet pretraining inflated to 3D, we isolate the token-mixing mechanism by comparing two primary configurations: Mod A (Mamba encoder with a CNN decoder) and Mod B (CNN encoder with a Mamba decoder). We benchmarked these configurations against established models, including nnU-Net, TransUNet, and Swin UNETR. Our findings indicate that integrating Mamba within the encoder (Mod A) consistently yields the strongest performance, achieving a Whole Tumour (WT) Dice score of 0.873 and a 95% Hausdorff Distance (HD95) of 12.39 mm. This approaches the full-resolution accuracy of nnU-Net (WT Dice 0.908). Conversely, deploying Mamba in the decoder (Mod B) substantially degrades spatial reconstruction (WT HD95 of 29.42 mm). Attention-heavy baselines like Swin UNETR severely underperformed (WT Dice 0.478) on this heterogeneous dataset. Our contributions are threefold: 1. Systematic Ablation: A controlled evaluation isolating Mamba's placement in the encoder versus decoder. 2. Clinical Benchmarking: Robustness assessment on the highly heterogeneous BraTS-Africa 2025 dataset. 3. Architectural Guidance: Evidence that Mamba excels in feature encoding, while CNNs retain an advantage in spatial decoding. Ultimately, our results demonstrate that token mixer placement is absolutely critical. Mamba excels at global encoding, whereas CNN decoders better reconstruct irregular tumour boundaries. This highly scalable hybrid blueprint is ideal for resource-constrained healthcare systems.