A Mechanistic Interpretability Analysis of Medical Modality Distractions in CT and MRI Abdominal Segmentation with SAM
Abstract
The Segment Anything Model (SAM) transfers unevenly to medical imaging, with performance depending strongly on the acquisition modality. We ask a mechanistic question: does SAM’s image encoder carry an explicit internal representation of imaging modality, and does it causally affect segmentation? We benchmark SAM ViT-B on CT versus MRI across three abdominal organs (liver, spleen, left kidney) from AMOS22 and CHAOS, holding anatomy, pixel spacing, orientation, and the box prompt fixed. MRI outperforms CT in every organ, with a Dice gap of 0.201 on CHAOS liver. We train a sparse autoencoder (SAE) on the 768-dimensional residual stream leaving the last ViT block (block 11), the tensor the neck projection consumes before the mask decoder sees anything, and identify 260 CT-ness and 71 MRI-ness latents that fire on one modality regardless of organ. These latents are structural rather than brightness-driven, transfer to unseen scanners at 79–93%, and occupy opposite spatial geometries: CT-ness fires outside the organ on the body wall and scanner table, while MRI-ness concentrates on tissue boundaries. Suppressing both modality-specific subspaces (271 latents that transfer to held-out CHAOS scanners) via delta patching at α= 0.25 closes 44% of the CT–MRI gap. Suppressing MRI-ness alone does nothing. We repeat the entire experiment for MedSAM, whose block-11 dictionary is even more modality-dominated: 6,107 of 12,288 latents (50%) have a significant CT–MRI gap, and its 237 CT-ness and 93 MRI-ness latents carry higher separation (median AUROC 0.986 vs. 0.977) and larger magnitudes (top CT magnitude 7.09 vs. 1.85) than SAM’s. Yet suppressing these stronger features does not change Dice for either modality: MedSAM’s mask decoder, the only component fine-tuned, has learnt to ignore modality-specific distractions. This dissociation between representation and behavior confirms that SAM’s inconsistent CT performance is causally linked to modality-specific features the decoder has not learnt to suppress.