Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)

๐Ÿ“… 2025-07-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of improving prediction accuracy for distributed cortical fMRI responses evoked by naturalistic film stimuli. To model multimodal temporal dynamics, we propose a hierarchical multimodal recurrent ensemble model: modality-specific bidirectional RNNs encode temporal evolution of video, audio, and pretrained language embeddings; a hierarchical feature fusion mechanism is coupled with a curriculum learning strategy progressing from sensory to association cortices; and a lightweight subject-specific output head is trained with an MSEโ€“correlation composite loss. Robustness is enhanced via ensemble averaging over 100 model variants. Our approach ranked third in the Algonauts 2025 Challenge (overall r = 0.2094), achieved a peak average correlation of 0.63 across individual brain regions, and notably improved prediction performance for the most challenging subject (Subject 5).

Technology Category

Application Category

๐Ÿ“ Abstract
Accurately predicting distributed cortical responses to naturalistic stimuli requires models that integrate visual, auditory and semantic information over time. We present a hierarchical multimodal recurrent ensemble that maps pretrained video, audio, and language embeddings to fMRI time series recorded while four subjects watched almost 80 hours of movies provided by the Algonauts 2025 challenge. Modality-specific bidirectional RNNs encode temporal dynamics; their hidden states are fused and passed to a second recurrent layer, and lightweight subject-specific heads output responses for 1000 cortical parcels. Training relies on a composite MSE-correlation loss and a curriculum that gradually shifts emphasis from early sensory to late association regions. Averaging 100 model variants further boosts robustness. The resulting system ranked third on the competition leaderboard, achieving an overall Pearson r = 0.2094 and the highest single-parcel peak score (mean r = 0.63) among all participants, with particularly strong gains for the most challenging subject (Subject 5). The approach establishes a simple, extensible baseline for future multimodal brain-encoding benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Predict cortical responses to naturalistic movies using multimodal integration
Map video, audio, and language embeddings to fMRI time series
Improve robustness and accuracy in brain-encoding benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical multimodal recurrent ensemble model
Bidirectional RNNs for temporal dynamics encoding
Composite MSE-correlation loss with curriculum training
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Semih Eren
Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; TU Dresden, Dresden, Germany
D
Deniz Kucukahmetler
Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; School for Embedded and Composite AI (SECAI), Dresden/Leipzig, Germany
Nico Scherf
Nico Scherf
Max Planck Institute for Human Cognitive and Brain Sciences, SCADS.AI, Leipzig University
Machine LearningComputational StatisticsData VisualizationArtificial Intelligence