邁向量化 ASR 模型中的基準最佳化
Towards Quantifying Benchmark Optimization in ASR Models
August 20, 2026
作者: Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis
cs.AI
摘要
公開基準是衡量自動語音辨識(ASR)模型能力的重要指標。然而,由於這些基準本質上具有公開性,模型存在針對這些基準進行最佳化的風險,而此類最佳化可能無法有效泛化至真實世界資料。我們提出一套量化基準最佳化的方法論,聚焦於音訊對參考轉錄文本呈現欠定性(underdetermined)的情境。我們識別出三類行為探針(behavioral probes),可用以揭示模型在音訊欠定的情況下再現基準參考片段的能力:參考不一致(reference disagreement)、遮罩數字還原(masked-number recovery),以及正字法切換(orthographic switching)。我們發現,得分最高的開源模型即使在相關音訊出現矛盾、被遮罩或模糊不清時,仍會輸出與基準參考轉錄片段完全一致的內容。透過多種機制性探針,我們證明模型會回應狹義的聲學線索,以覆蓋對音訊的忠實表徵,轉而採用基準最佳化策略。我們進一步顯示,此基準最佳化行為可透過低秩線性引導(low-rank linear steering)進行因果性操控,在某些情況下甚至只需在片段末尾附加音訊即可觸發。整體而言,我們的結果表明,高分模型表現出受基準條件制約的行為,此類行為可能虛增基準效能,卻未必反映通用轉錄能力的實質提升。
English
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.