Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offs

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过比较四种架构策略,解决了流式多说话人语音识别中准确性、延迟和效率的平衡问题。
📝 Abstract
Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.
Problem

Research questions and friction points this paper is trying to address.

Streaming multi-speaker ASR
accuracy
latency
efficiency
overlapping speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Multi-Speaker ASR
Architectural Strategies
Diarization Integration
Systematic Analysis
🔎 Similar Papers
No similar papers found.