ChatPaper.aiChatPaper

ASRモデルにおけるベンチマーク最適化の定量化に向けて

Towards Quantifying Benchmark Optimization in ASR Models

August 20, 2026
著者: Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis
cs.AI

要旨

公開ベンチマークは、自動音声認識(ASR)モデルの能力を測る重要な指標である。しかし、公開されているという性質上、モデルがこれらのベンチマークに対して最適化され、実世界のデータにはうまく一般化しないというリスクが存在する。本稿では、音声が参照書き起こしを過少決定するケースに焦点を当て、ベンチマーク最適化を定量化する方法論を提示する。参照不一致(reference disagreement)、マスク数字の復元(masked-number recovery)、表記法切り替え(orthographic switching)という、過少決定された音声にもかかわらずベンチマーク参照スパンを再現するモデルの能力を明らかにする3つの行動プローブ群を特定する。最高スコアのオープンソースモデルは、関連する音声が矛盾している、マスクされている、あるいは曖昧である場合でも、ベンチマーク参照書き起こしスパンを逐語的に出力することがわかった。様々な機構的プローブを用いることで、モデルが限定的な音響的手がかりに反応し、音声の忠実な表現を覆して、ベンチマーク最適化された方策を優先することを示す。さらに、ベンチマーク最適化された行動が、低ランク線形ステアリングや、場合によってはセグメントの末尾に音声を単純に追加することによって、因果的に操作できることを示す。全体として、我々の結果は、高性能モデルがベンチマーク条件付きの行動を示し、その行動が改善された汎用転写能力を反映することなくベンチマーク性能を膨張させ得ることを示している。
English
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.