Multi-modal Integration Analysis of Alzheimer's Disease Using Large Language Models and Knowledge Graphs
Alzheimer’s disease (AD) research faces challenges in integrating heterogeneous, unpaired, and cross-cohort multimodal data—including MRI, gene expression, biomarkers, EEG, and clinical metrics—due to their distributed nature and lack of subject-level alignment. Method: We propose the first large language model (LLM)-driven knowledge graph reasoning framework enabling population-level, ID-agnostic, concept-level cross-modal association mining and natural language hypothesis generation. Our approach integrates multimodal statistical feature selection, cross-cohort cross-validation, and expert consensus evaluation (Cohen’s κ = 0.82). Contribution/Results: We identify a novel pathological cascade—“metabolic risk → neuroinflammation → tau dysregulation”—and robust frontal EEG–gene expression associations (r = 0.42–0.58, p < 0.01; high-significance links: r > 0.6, p < 0.001), with effect sizes stable across cohorts (variance < 15%). These findings yield testable, mechanistically grounded hypotheses for AD pathogenesis and therapeutic target discovery.