🤖 AI Summary
This study addresses the limitation of existing deep learning models in neuroimaging, which are typically confined to single tasks and struggle with cross-task knowledge transfer. To overcome this, the authors propose GenFAR, a modular deep learning framework trained jointly on 17 cognitive, clinical, and diagnostic tasks using brain MRI data from 49,246 individuals across 11 cohorts. The framework leverages an innovative task-sequencing mechanism and a novel Donor Score metric to identify five pivotal source tasks that substantially enhance sample efficiency and performance on downstream tasks. The resulting general-purpose brain representations demonstrate strong clinical relevance, significantly improving model accuracy on unseen tasks while markedly reducing the required training sample size.
📝 Abstract
Deep learning models for neuroimaging have largely been developed for individual tasks, limiting knowledge transfer across applications. Here we introduce GenFAR, a modular deep learning framework that learns general, clinically informed features from brain MRIs. We trained this modular architecture on 49,246 individuals across 11 cohorts, using 17 diverse classification and regression tasks spanning cognition, clinical, diagnosis, demographics, and biomarkers. This yields aggregated, focused feature sets that capture rich, clinically- and biologically-relevant brain representations. We developed a sequential learning approach where tasks progressively build on previously learned representations. Through an analysis of 5,000 task sequences, we identified an optimal sequence length of six tasks and introduced a Donor Score metric to quantify each task's contribution to downstream performance. This analysis revealed five consistently strong donor tasks (Age, AD/MCI, MMSE, Hypertension, Hyperlipidemia) that formed the base of our sequential model. We demonstrated the utility of our learned representation, in various tasks beyond those included in the training set, to serve as the foundation for specialized secondary predictors. We further showed that using the learned feature representation can substantially increase the sample efficiency of secondary deep learning training tasks and models, as well as improve their accuracy.