When Audio Residuals Are Still Visible: A Faithfulness Audit for Foundation-Model-Based fMRI Decoding
Abstract
Foundation models are used for the interpretation of scientific data, but whether those interpretations are accurate requires an evaluation that isolates the specific factor that is to be analyzed. In this research, we study cross-modal fMRI decoding using naturalistic audiovisual movies and ImageBind representations. We trained a neural decoder to only learn from the video representations of the movies, with no audio supervision. The audio representations were separately residualized with respect to video before we tested whether the brain-derived predictions could retrieve those residualized audio targets. Under the originally specified ridge residualizer, true video alone retrieved the residualized audio targets at 7.44× chance. Under additional residualization and robustness analyses, true-video retrieval remained above chance, including 3.74× under validation-selected ridge and 4.58× under a nonlinear residualizer. The brain-derived predictions reached 2.77× chance. A validation-calibrated video-mixture control reached 2.04× chance, while a Gaussian control reached 1.49× chance. This does not indicate that the brain is producing auditory interpretations of videos. Instead, it shows that the same type of retrieval outcome remains accessible from true video even when no brain data are used. Thus, the brain-derived retrieval effect cannot be uniquely interpreted as evidence of audio-related neural information beyond the concurrent visual stimulus. In general, when foundation models are applied to science, it is critical to ensure that transformed targets are directly tested for their accessibility through nuisance routes, under the same representation and evaluation metric used for the scientific interpretation.