ChatPaper.aiChatPaper

ASR 모델의 벤치마크 최적화 정량화를 향하여

Towards Quantifying Benchmark Optimization in ASR Models

August 20, 2026
저자: Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis
cs.AI

초록

공개 벤치마크는 자동 음성 인식(ASR) 모델의 성능을 측정하는 중요한 지표이다. 그러나 공개 데이터라는 특성상, 모델이 실제 세계 데이터에는 제대로 일반화되지 않는 방식으로 이러한 벤치마크에 맞춰 최적화될 위험이 존재한다. 본 연구는 오디오가 참조 전사(reference transcript)를 충분히 결정하지 못하는 경우에 초점을 맞추어, 벤치마크 최적화를 정량화하는 방법론을 제시한다. 우리는 오디오가 불충분하게 결정된 상황에서도 모델이 벤치마크 참조 구간을 재현하는 능력을 드러내는 세 가지 행동 탐침(behavioral probe)군을 식별한다: 참조 불일치(reference disagreement), 마스킹된 숫자 복원(masked-number recovery), 그리고 철자 전환(orthographic switching). 우리는 최고 점수를 기록한 오픈소스 모델들이 관련 오디오가 모순되거나, 마스킹되거나, 모호한 경우에도 벤치마크 참조 전사 구간을 문자 그대로 출력함을 발견한다. 다양한 기계적 탐침(mechanistic probe)을 통해, 모델이 좁은 음향적 단서에 반응하여 오디오의 충실한 표현을 무시하고 벤치마크 최적화 정책을 우선시함을 보여준다. 또한 이러한 벤치마크 최적화 행동이 저랭크 선형 조향(low-rank linear steering)이나 일부 경우에는 단순히 세그먼트 끝에 오디오를 추가하는 방식으로 인과적으로 조작될 수 있음을 입증한다. 종합하면, 본 연구의 결과는 고성능 모델들이 벤치마크 조건화된 행동을 나타내며, 이는 실제 범용 전사 능력의 향상을 반영하지 않으면서 벤치마크 성능을 부풀릴 수 있음을 시사한다.
English
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.