REFORM-3D: A Representation-Centric Evaluation Framework for 3D Medical Vision Foundation Models
Abstract
Evaluation of 3D medical foundation models is still dominated by downstream task scores, which entangle representation quality with adaptation choices. This is especially critical in medical imaging, where spacing and anisotropy can substantially alter image statistics while the anatomy itself remains unchanged. We introduce \textbf{REFORM-3D}, a \textbf{R}epresentation-centric \textbf{E}valuation \textbf{F}ramew\textbf{OR}k for benchmarking frozen 3D medical vision foundation \textbf{M}odels across three questions: Acquisition Robustness, asking whether features remain stable when spacing varies while anatomy is fixed; Cross-Modal Anatomical Alignment, asking whether same-organ CT-MR pairs remain recoverable under balanced bidirectional retrieval; and Anatomical Generalization, asking whether limited adaptation recovers held-out organ signals beyond a frozen-feature baseline under case-disjoint evaluation. These are also pretraining questions, because what a 3D encoder learns depends on how volumes are geometrically presented during self-supervised pretraining. We therefore employ REFORM-3D to evaluate encoders pretrained on 62K unlabeled multimodal 3D volumes spanning diverse anatomical structures across MR, CT, PET, CTA, and CBCT under three explicit geometric assumptions: isotropic sampling, voxel-relative sampling, and spacing-aware resampling. By exposing the encoder to acquisition geometry during pretraining rather than treating geometry only as a downstream issue, REFORM-3D evaluates whether learned representations preserve stable anatomical structures or remain sensitive to acquisition and modality shortcuts.