Towards Quantifying Benchmark Optimization in ASR Models

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种量化自动语音识别模型对公开基准过拟合的方法,通过设计行为探针揭示了模型在音频不确定情况下仍能复现参考文本的问题。
📝 Abstract
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Benchmark Optimization
Generalization
Real-World Data
Model Overfitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark Optimization
Behavioral Probes
Low-rank Linear Steering
🔎 Similar Papers
No similar papers found.