HearInContext: A Benchmark for Implicit Context in Speech Recognition

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决语音识别中的隐式上下文问题,本文引入了HearInContext基准测试,通过同音词构建测试案例,并使用Qwen3-ASR-1.7B模型微调提高了目标召回率。
📝 Abstract
Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context. We introduce HearInContext, a Mandarin--English benchmark that pairs shared synthetic speech with assistant replies supporting different interpretations. The benchmark comprises 3,764 semantic test cases built around homophones. Implicit contexts exclude candidate words; explicit contexts name the target. No-context and unrelated-context controls measure the benefit of relevant history and sensitivity to irrelevant history. Context-capable models benefit from implicit cues but achieve higher target recall with explicit hints. Fine-tuning Qwen3-ASR-1.7B improves implicit-context target recall by 11.0 and 11.5 percentage points in Mandarin and English, respectively, while absolute CER/WER changes on AISHELL-1 and LibriSpeech remain below 0.1 percentage points. Gains extend to explicit conditions excluded from fine-tuning and to Mandarin hotword recognition on real recordings.
Problem

Research questions and friction points this paper is trying to address.

Speech Recognition
Implicit Context
Homophones
Innovation

Methods, ideas, or system contributions that make the work stand out.

HearInContext
implicit context
speech recognition
homophones
contextual ASR
🔎 Similar Papers
No similar papers found.
Y
Yifan Gao
AI Center, OPPO, Beijing, China
Y
Yao Tian
AI Center, OPPO, Beijing, China
H
Hongbin Suo
AI Center, OPPO, Beijing, China