StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses performance degradation caused by label scarcity in domain adaptation for streaming automatic speech recognition (ASR). We propose a semi-supervised streaming ASR framework that fine-tunes a student model using pseudo-labels generated by an offline transducer teacher. Crucially, we introduce a novel dynamic programming realignment mechanism with ASR anchor-based prior regularization to effectively correct chunk-level token misalignments and mitigate domain shift. Experimental results across four datasets demonstrate that this framework consistently outperforms supervised fine-tuning baselines and significantly narrows the performance gap with the offline teacher model. These findings validate the effectiveness of our approach in low-resource cross-domain scenarios, offering a robust solution for adapting streaming ASR systems where annotated data is limited.
📝 Abstract
Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that adapts a pretrained streaming student by fine-tuning an offline transducer teacher on the labeled training set, generating pseudo-labels on the unlabeled portion, and fine-tuning the student on the mixture. We further introduce a prior-regularized dynamic-programming realignment step that fixes chunk-level word placement using an ASR-hypothesis anchor. Across four datasets spanning financial calls, prepared read speech, and phone-quality dialogue, StreamHear consistently outperforms supervised student fine-tuning and narrows the gap to the offline teacher.
Problem

Research questions and friction points this paper is trying to address.

Streaming ASR
Domain Shift
Semi-Supervised Learning
Pseudo-Labeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semi-Supervised Learning
Pseudo-Labeling
Domain Adaptation
Streaming ASR
Dynamic-Programming Realignment
🔎 Similar Papers
No similar papers found.
Z
Zefang Liu
Capital One, USA
C
Chenyang Zhu
Capital One, USA
Sangwoo Cho
Sangwoo Cho
Capital One
Natural Language ProcessingComputer VisionDeep LearningMachine Learning
X
Xujun Peng
Capital One, USA
S
Shi-Xiong Zhang
Capital One, USA
Sambit Sahu
Sambit Sahu
Capital One
Generative AILLM Pre-trainingInference Optimization