Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有AVSR基准数据集缺乏自然对话复杂性的问题,通过构建Candor-LR数据集,利用更真实的对话场景来提高模型性能和跨域鲁棒性。
📝 Abstract
Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field toward realistic dialogue, we introduce Candor-LR, a conversational benchmark derived from the CANDOR corpus of 1,656 natural dyadic videoconferences. Our custom data preparation pipeline yields 713.5, 10.1, and 60.1 hours of training, validation, and test data, respectively. Evaluating pretrained AVSR models on Candor-LR reveals that audio-only accuracy drops sharply compared to LRS3, but visual cues compensate effectively, driving much larger performance gains on Candor-LR than on LRS3. Furthermore, training on this corpus significantly improves cross-domain robustness under both clean and noisy conditions, as its realistic conversational data captures broader audio-video features. We open-source our pipeline to ensure reproducibility, establishing Candor-LR as a challenging benchmark for conversational AVSR.
Problem

Research questions and friction points this paper is trying to address.

audio-visual speech recognition
natural conversation
realistic dialogue
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dyadic Conversational Dataset
Audio-Visual Speech Recognition
Realistic Dialogue
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13
💼 Related Jobs
No related jobs found.
R
Rishabh Jain
Sigmedia Group, School of Engineering, Trinity College Dublin, Ireland
A
Aristeidis Papadopoulos
Sigmedia Group, School of Engineering, Trinity College Dublin, Ireland
Zhaofeng Lin
Zhaofeng Lin
Trinity College Dublin
audio-visual speech recognitionautomatic speech recognition
Naomi Harte
Naomi Harte
Professor in Speech Technology, Trinity College Dublin
Audio-visual speech recognitionspeech qualitymultimodal interactionbirdsong analysis