The Brain's Director's Cut: Neural Reranking for fMRI-to-Video Reconstruction
Abstract
Visual brain decoding from functional magnetic resonance imaging (fMRI) data has made remarkable progress in recent years, driven by increasingly powerful artificial intelligence (AI) models and the growing availability of large-scale public datasets. However, despite the impressive advances achieved in static image decoding, the reconstruction of videos from fMRI signals remains challenging due to the temporal mismatch between rapidly changing video stimuli and the slow dynamics of the blood-level-oxygen-dependent (BOLD) response. In this work, we show on the publicly available CC2017, a benchmark dataset for video-fMRI experiments, that a simple Bayesian-inspired resampling strategy can achieve performance comparable to the state-of-the-art (SOTA) while estimating only video text-conditioning using linear models, instead of relying on complex multi-stage pipelines, commonly needed in the video decoding literature. Experiments across three subjects show that our approach consistently benefits the semantic quality of the reconstructed videos. These findings suggest an alternative and simpler strategy for improving video brain decoding, highlighting the potential of our implementation as a complementary direction to increasingly complex pipelines. The code is available at https://github.com/nessuno-source/video_decoding.