ChatPaper.aiChatPaper

面向ASR模型中基准优化的量化研究

Towards Quantifying Benchmark Optimization in ASR Models

August 20, 2026
作者: Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis
cs.AI

摘要

公共基准是自动语音识别(ASR)模型能力的重要衡量标准。然而,由于其公开性质,模型存在以不能很好地泛化到真实世界数据的方式针对这些基准进行优化的风险。我们提出了一种量化基准优化的方法论,重点关注音频对参考转录文本欠定的情况。我们确定了三类行为探针,它们揭示了模型在音频欠定的情况下复现基准参考片段的能力:参考标注不一致、掩码数字恢复和正字法切换。我们发现,得分最高的开源模型即使在相关音频是矛盾的、被掩码的或模糊的情况下,也会输出逐字的参考转录片段。通过使用多种机制探针,我们表明模型会响应狭窄的声学线索,以覆盖对音频的忠实表征,转而采用基准优化的策略。我们表明,这种基准优化行为可以通过低秩线性引导进行因果操纵,或者在某些情况下,只需在片段末尾追加音频即可触发。总体而言,我们的结果表明,高性能模型表现出基准条件化行为,这些行为可能虚增基准性能,而并不反映通用转录能力的提升。
English
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.