Towards Quantifying Benchmark Optimization in ASR Models
本文提出了一种量化自动语音识别模型对公开基准过拟合的方法,通过设计行为探针揭示了模型在音频不确定情况下仍能复现参考文本的问题。
本文提出了一种量化自动语音识别模型对公开基准过拟合的方法,通过设计行为探针揭示了模型在音频不确定情况下仍能复现参考文本的问题。
Current evaluations of spoken AI systems often rely on isolated metrics—such as recognition accuracy or textual quality—overlooking the acoustic characteristics and multidimensional nature inherent to speech. This work proposes the first comprehensive evaluation framework that integrates acoustic fidelity, expressiveness, interactivity, and robustness, establishing a real-world benchmark spanning text-to-speech (TTS), speech-to-speech translation (STS), spoken understanding (SU), and automatic speech recognition (ASR). The framework employs diverse data encompassing accents, emotions, background noise, and conversational contexts for fine-grained assessment. Findings reveal that system performance is highly dimension-dependent: TTS exhibits relative independence across dimensions, STS frequently neglects acoustic expressivity, SU shows inconsistent performance on paralinguistic tasks, and ASR exposes critical weaknesses under realistic conditions that conventional benchmarks fail to capture—thereby underscoring the inadequacy of single-metric evaluations.
本文提出了一种量化自动语音识别模型对公开基准过拟合的方法,通过设计行为探针揭示了模型在音频不确定情况下仍能复现参考文本的问题。
Current evaluations of spoken AI systems often rely on isolated metrics—such as recognition accuracy or textual quality—overlooking the acoustic characteristics and multidimensional nature inherent to speech. This work proposes the first comprehensive evaluation framework that integrates acoustic fidelity, expressiveness, interactivity, and robustness, establishing a real-world benchmark spanning text-to-speech (TTS), speech-to-speech translation (STS), spoken understanding (SU), and automatic speech recognition (ASR). The framework employs diverse data encompassing accents, emotions, background noise, and conversational contexts for fine-grained assessment. Findings reveal that system performance is highly dimension-dependent: TTS exhibits relative independence across dimensions, STS frequently neglects acoustic expressivity, SU shows inconsistent performance on paralinguistic tasks, and ASR exposes critical weaknesses under realistic conditions that conventional benchmarks fail to capture—thereby underscoring the inadequacy of single-metric evaluations.