TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TurboBias 2.0,通过GPU加速和批处理解码技术,解决了生产环境下语音识别系统中实时、高效地使用个性化上下文偏置的问题。
📝 Abstract
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, they often do not address the practical requirements of modern production ASR systems: streaming inference, efficient batched decoding, user-specific context lists, and low runtime overhead. We propose TurboBias 2.0, a production-oriented framework for efficient phrase boosting in Transducer-based ASR systems. The framework extends GPU-accelerated TurboBias with a case-insensitive boosting graph and per-stream batched decoding, allowing each utterance in a batch to use an independent context-biasing configuration. This enables personalized context biasing for multiple simultaneous users without sharing or mixing their context lists. The proposed framework supports both offline and streaming inference and can be used with greedy and beam-search decoding. Experiments show that TurboBias 2.0 improves contextual phrase recognition while preserving low latency and high throughput.
Problem

Research questions and friction points this paper is trying to address.

context-biasing
production ASR systems
streaming inference
batched decoding
user-specific context lists
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Context-Biasing
Production-Efficient ASR
TurboBias 2.0
User-Specific Context Lists
Batched Decoding
🔎 Similar Papers
No similar papers found.