DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of source language selection in zero-shot cross-lingual speech recognition for low-resource languages, where performance is hindered by linguistic divergence, orthographic inconsistency, and uneven resource distribution. The authors propose DonorRank, the first framework to apply learning-to-rank to this task, integrating multilingual speech representations with linguistic metrics to predict optimal source language combinations. Experimental results on Indic and African language datasets demonstrate that DonorRank significantly outperforms heuristic strategies based on phylogenetic similarity or reliance on high-resource languages. Beyond accurately ranking effective donor languages, the approach reveals the critical influence of source language set composition on transfer performance, offering both interpretable insights and practical guidance for low-resource automatic speech recognition.
📝 Abstract
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.
Problem

Research questions and friction points this paper is trying to address.

donor language selection
low-resource ASR
cross-lingual transfer
spontaneous speech
multilingual speech recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

DonorRank
cross-lingual transfer
low-resource ASR
learning-to-rank
donor language selection
💼 Related Jobs
No related jobs found.
A
Akriti Dhasmana
Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, USA
A
Aarohi Srivastava
Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, USA
David Chiang
David Chiang
Associate Professor, University of Notre Dame
Natural Language ProcessingMachine Translation