Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

📅 2026-05-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing neurolinguistic dementia detection approaches, which are predominantly English-centric and ill-suited for low-resource multilingual clinical settings involving Tagalog–English code-switching. The authors construct the first Tagalog–English parallel dementia dialogue dataset comprising 4,000 human-translated utterances and systematically evaluate models—including TF-IDF with logistic regression, BERT, NeoBERT, XLM-R, and RoBERTa-Tagalog—across monolingual, zero-shot cross-lingual, and bilingual fine-tuning setups. Notably, NeoBERT is introduced for the first time in clinical NLP, and the work establishes the first systematic benchmark for dementia detection in Tagalog. Experimental results demonstrate that bilingual fine-tuning effectively mitigates cross-lingual performance degradation, enabling all Transformer-based models to achieve Macro-F1 scores of 0.969–0.973 on Tagalog, thereby underscoring that linguistic coverage outweighs model size or architecture in determining performance in multilingual clinical NLP.
📝 Abstract
Dementia detection from spontaneous speech offers a scalable approach to cognitive screening, yet NLP systems remain predominantly English-centric. This limitation is especially acute in the Philippines, where Filipino-English code-switching is pervasive and no prior work has addressed NLP-based dementia detection. We present the first systematic evaluation of transformer-based dementia detection in Filipino speech and the first assessment of NeoBERT in a clinical NLP setting. To separate language from domain effects, we construct a parallel bilingual dataset of 4,000 DementiaBank-derived transcripts, with Filipino translations produced manually to preserve discourse-level markers of cognitive decline. We evaluate five model families, TF-IDF + LogReg, BERT, NeoBERT, XLM-R, and RoBERTa-Tagalog, under monolingual, zero-shot cross-lingual, and bilingual fine-tuning settings. We find that in-domain performance does not transfer across languages, with English-trained BERT dropping to Macro-F1 = 0.455 on Filipino, and that architectural modernization alone does not improve robustness. Bilingual fine-tuning, however, eliminates cross-lingual degradation across all transformer models, converging to Macro-F1 = 0.969-0.973. These results suggest that multilingual clinical NLP performance is driven primarily by linguistic coverage during training rather than model scale or architecture.
Problem

Research questions and friction points this paper is trying to address.

dementia detection
low-resource languages
Filipino-English code-switching
clinical NLP
cross-lingual transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

dementia detection
low-resource languages
code-switching
bilingual fine-tuning
NeoBERT
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Rez Samantha Z. Floresca
Ateneo de Manila Senior High School
E
Edric Castel C. Hao
Analog Devices, Inc.
H
Hannah Grachiella Buñales
Ateneo de Manila Senior High School
C
Chelsea Dominique E. Temprosa
Ateneo de Manila Senior High School
G
Georgianna Z. Reyes
Ateneo de Manila Senior High School
K
Kervin Gabriel L. Chua
Ateneo de Manila Senior High School