Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对韩语语音识别中的错误修正问题,提出了一种基于文本的后编辑框架DCSC,并引入了首个大规模韩语对话级ASR错误修正基准数据集DasanCallDial。
📝 Abstract
Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restrictions preclude audio access, error correction must rely on text-based post-editing. Existing text-only approaches face significant challenges in low-resource languages, mainly due to a critical scarcity of annotated corpora and tailored correction methodologies. For Korean, this resource gap is particularly pronounced, as existing resources are predominantly designed for ASR training rather than text-based error correction. To address this, we introduce DasanCallDial, the first large-scale Korean benchmark dataset specifically curated for dialogue-level ASR error correction. Derived from genuine call center interactions, it comprises 1,974 dialogues with 115,460 utterances. Leveraging this resource, we propose Detector-Gated Contextual Span Correction (DCSC), a text-only post-editing framework for error-sparse Korean speech recognition transcripts. DCSC combines an encoder-based detector that first performs token-level error detection, followed by a language model-based corrector trained to rectify fine-grained span-level errors. Additionally, we employ dialogue-level context augmentation to enable the model to leverage discourse history for disambiguation. By employing multi-level granularity, our method achieves state-of-the-art performance, effectively overcoming the limitations of general LLMs in low-resource settings.
Problem

Research questions and friction points this paper is trying to address.

Korean Speech Recognition
Error Correction
Consultation Services
Text-based Post-editing
Low-resource Languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Detector-Gated Contextual Span Correction
text-only post-editing
low-resource languages
dialogue-level context augmentation
Korean speech recognition
🔎 Similar Papers
No similar papers found.
Y
Yonghyun Jun
Chung-Ang University
J
Jimin Lee
Chung-Ang University
H
Hwan Chang
Chung-Ang University
D
Dongho Shin
Korea Local Information Research & Development Institute
S
Seolah Kim
SK intellix
Hwanhee Lee
Hwanhee Lee
Assistant Professor, Department of Artificial Intelligence, Chung-Ang University
Natural Language ProcessingTrustworthy LLMLLM Safety