AVSRBench: A Multi-Condition AVSR Benchmark

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在多种条件下评估三种AVSR架构,揭示了现有视觉语音识别系统在广播领域外的泛化能力不足,并引入RoomReader-AV作为新基准以解决此问题。
📝 Abstract
While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipspeakers and non-professional speakers, and spontaneous multi-party video conversations. We find that visual-only performance deteriorates rapidly beyond broadcast domains, and audio-video fusion mainly benefits Lombard speech environments. Visual understanding degrades sharply at 90° profile views, with multimodal systems relying largely on acoustic fallback. Additionally, speaker articulation proves more critical than minor camera shifts, and LLM-based architectures suffer from poor out-of-domain generalization. Our work highlights a significant generalization gap in current AVSR research. To address this, we also introduce RoomReader-AV as a new benchmark for AVSR and release a unified data preprocessing pipeline to make comprehensive multi-condition evaluation accessible.
Problem

Research questions and friction points this paper is trying to address.

AVSR
generalization
benchmark
broadcast speech
domain adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-condition benchmark
generalization gap
audio-video fusion
Lombard speech
unified preprocessing pipeline
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.