Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了语音预处理和数据集整理对阿尔茨海默病检测的影响,通过对比不同处理方法下的模型泛化能力,发现增强语音虽然提高了域内性能,但降低了跨域鲁棒性。
📝 Abstract
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
dataset curation
Alzheimer's disease detection
model generalization
cross-domain evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

speech enhancement
cross-dataset generalization
large audio-language models (LALMs)
class imbalance
prediction shifts
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
L
Luqi Sun
Center for Language and Speech Processing (CLSP), Johns Hopkins University
S
Shreeram Suresh Chandra
Center for Language and Speech Processing (CLSP), Johns Hopkins University
L
Lin Zhang
Center for Language and Speech Processing (CLSP), Johns Hopkins University
Y
You-Jin Li
Research Center for Information Technology Innovation, Academia Sinica
Brian MacWhinney
Brian MacWhinney
University of Michigan, Ann Arbor
Yu Tsao
Yu Tsao
Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment
Emily Mower Provost
Emily Mower Provost
Professor of Computer Science, University of Michigan
Emotion RecognitionMachine LearningEmotion Perception
Berrak Sisman
Berrak Sisman
Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice ConversionAnti-spoofing