MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
Abstract
Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been driven largely by public datasets focused on brain and knee imaging, shaping how models are trained and evaluated. As a result, careful studies of the reliability of these models across diverse anatomical settings remain limited. In this work, we introduce MosaicMRI, a large and diverse collection of fully sampled raw musculoskeletal (MSK) MR measurements designed for training and evaluating machine-learning--based methods. MosaicMRI is the largest open-source raw MSK MRI dataset to date, comprising 2,725 volumes and 81,570 slices, including a 1.5T core dataset and a 0.55T low-field set. The dataset offers substantial diversity in volume orientation (e.g., axial, sagittal), imaging contrasts (e.g., PD, T1, T2), anatomies (e.g., spine, knee, hip, ankle, and others), numbers of acquisition coils, and field strengths. Using accelerated reconstruction as a testbed, we study scaling, robustness, and data selection under realistic MSK distribution shifts. Across E2E-VarNet, U-Net, and ViT baselines, mixed-anatomy training improves over anatomy-specific training, showing that the benefit of anatomical diversity is not architecture-specific. Zero-shot experiments on the 0.55~T low-field set highlight that using more training data improves average reconstruction under field-strength shift. Controlled subset and leave-one-anatomy-out experiments further suggest that these gains come from transferable cross-anatomy structure, rather than anatomy or contrast matching alone.