Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多说话人语音识别中的说话人泄漏问题,提出了一种基于说话人日志的剪枝方法来修正转录文本,提高了复杂声学环境下转录的可靠性。
📝 Abstract
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identifies and removes leakage artifacts. Our method utilizes a pre-trained speaker diarization model as a multimodal verifier to prune transcribed segments satisfying a tripartite consensus of temporal containment, lexical cross-validation, and temporal alignment. Results on LibriMix, LibriSpeechMix, and the AMI Meeting corpus show our algorithm consistently reduces cpW ER across diverse overlap conditions. Specifically, on subsets with high speaker leakage, our method achieves relative cpW ER reductions of up to 29%, highlighting its effectiveness in enhancing the reliability of cascaded MT-ASR transcripts in complex acoustic environments.
Problem

Research questions and friction points this paper is trying to address.

Speaker Leakage
Cascaded Multi-talker ASR
Diarization
Innovation

Methods, ideas, or system contributions that make the work stand out.

speaker leakage
pruning-based paradigm
diarization model
tripartite consensus
cascaded MT-ASR
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hermann Yepdjio Nkouanga
Portland State University, USA
M
Minwei Luo
Portland State University, USA
Maggie Wigness
Maggie Wigness
US Army Research Laboratory
Computer VisionMachine LearningRobotics
Suresh Singh
Suresh Singh
Portland State University, USA