🤖 AI Summary
This study addresses the critical tension between privacy preservation and diagnostic utility in voice-based dementia detection, where speaker identity is often inadvertently leaked. To reconcile this trade-off, the authors propose the first multi-level framework that jointly optimizes privacy protection at both signal and feature levels. At the signal level, cumulative signal attack (CSA) precisely perturbs keyword regions; at the feature level, a gradient reversal layer (GRL) combined with mutual information-guided noise injection enables de-identified representation learning. Evaluated on the DementiaBank Pitt Corpus, the method achieves strong privacy guarantees—evidenced by a speaker verification equal error rate (EER) of 0.59 and near-zero speaker identification F1 score—while maintaining high dementia classification performance (F1 = 0.78, AUC = 0.86).
📝 Abstract
Speech recordings used for dementia detection inherently expose speaker identity, raising critical privacy concerns. Existing methods typically address only singular threats and fail to resolve the privacy--utility trade-off. We propose a multi-level framework designed to neutralize two distinct eavesdropping vectors. At the signal level, a Cumulative Signal Attack (CSA) concentrates perturbations in keyword-aligned regions to maximize transcription error (Word Error Rate WER = 1.00) while preserving vital prosodic biomarkers. At the feature level, a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection suppresses speaker-discriminative dimensions while retaining dementia-relevant diagnostic structure. Evaluated on the DementiaBank Pitt Corpus, our framework achieves near-chance speaker identification (Equal Error Rate EER = 0.59, F1 = 0.003) while maintaining strong dementia classification performance (F1 = 0.78, AUC = 0.86).