Many Brains, One Geometry: A Shared Visual-Semantic Space for Cross-Dataset fMRI Decoding
Moein Khajehnejad ⋅ Michelangelo Tronti ⋅ Forough Habibollahi ⋅ Tommaso Boccato ⋅ Matteo Ferrante ⋅ Nicola Toschi
Abstract
Visual-fMRI datasets vary widely in subjects, scanners, acquisition protocols, and stimuli, raising the question of whether they share a common representational geometry. We investigate this using MOSAIC, which combines eight datasets, 93 dataset-specific participant entries, and 430,007 single-trial fMRI–stimulus pairs. A subject-conditioned ROI-wise Transformer maps responses from 389 brain regions into the 512-dimensional CLIP space, using a multi-positive contrastive objective to align repeated stimuli across subjects and datasets. The model supports retrieval across seven evaluation datasets. On eight matched participant entries, the model achieves $35.0\pm11.1\%$ Top-10 accuracy, outperforming the main baselines—the MindEye-style pooled-CLIP decoder ($27.1\pm5.1\%$) and ridge regression ($21.1\pm9.1\%$)—and having the highest observed accuracy for seven of eight entries. With subject embeddings disabled, the encoder transfers zero-shot to unseen subjects and datasets, improving with broader pretraining by up to $92.1\%$. The learned brain space preserves graded semantic geometry across subjects and datasets, while ablations and saliency highlight ventral and early visual cortex and category-specific motion and attentional systems. These results support scalable cross-dataset decoding into a common CLIP-aligned space, with model sensitivity concentrated in ventral and early visual inputs.
Chat is not available.
Successful Page Load