Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
This work addresses the challenges of medical speech recognition, including domain-specific terminology, contextual ambiguity, and accurate transcription of clinical abbreviations and numerical values—issues that existing systems struggle to reconcile with real-time performance, accuracy, and generalization. The authors propose a modular decoupled architecture that separates the transcription pipeline into three specialized stages: domain-adapted recognition, formatting, and context-aware correction. This approach achieves, for the first time, high-recall recognition of medical terms and generates structured clinical text while supporting adaptive deployment across diverse scenarios. The system offers a production-grade API compatible with real-time dictation, conversational input, and batch processing. Evaluated on public medical speech datasets, it significantly outperforms state-of-the-art methods in the clinical domain while matching or exceeding their performance on general-domain tasks. The study also introduces the first Chinese clinical speech benchmark dataset to advance research in this area.