Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)
This study addresses the challenge of improving prediction accuracy for distributed cortical fMRI responses evoked by naturalistic film stimuli. To model multimodal temporal dynamics, we propose a hierarchical multimodal recurrent ensemble model: modality-specific bidirectional RNNs encode temporal evolution of video, audio, and pretrained language embeddings; a hierarchical feature fusion mechanism is coupled with a curriculum learning strategy progressing from sensory to association cortices; and a lightweight subject-specific output head is trained with an MSE–correlation composite loss. Robustness is enhanced via ensemble averaging over 100 model variants. Our approach ranked third in the Algonauts 2025 Challenge (overall r = 0.2094), achieved a peak average correlation of 0.63 across individual brain regions, and notably improved prediction performance for the most challenging subject (Subject 5).