LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening
This study addresses the critical gap in non-invasive, low-cost methods for early Alzheimer’s disease (AD) screening across diverse clinical settings. The authors propose a privacy-preserving approach that locally deploys open-source large language models to extract embeddings from automatically transcribed speech, followed by dimensionality reduction via principal component analysis (PCA) and machine learning–based classification—all without uploading sensitive patient data. This work represents the first integration of locally hosted large language models with transcribed speech for AD detection, achieving up to a 5% improvement in accuracy on the ADReSS20 and ADReSSo2021 datasets. Notably, the method demonstrates superior performance in the early stages of the disease and exhibits enhanced cross-dataset generalizability.