ChatPaper.aiChatPaper

잠재 변수로서의 전사 정책: 단어 수준 타이밍을 통한 제어 가능한 축어적 ASR 활성화

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

July 21, 2026
저자: Laurin Wagner, Mario Zusag, Bernhard Thallinger
cs.AI

초록

이질적으로 주석이 달린 데이터로 훈련된 최신 ASR 모델은 전사 스타일(축어적 대 의도적)을 통제되지 않은 잠재 변수로 취급하여, 측정 가능한 디코딩 불안정성, 평가 혼동(보고된 WER의 최대 60%가 스타일 불일치에 기인), 그리고 신뢰할 수 없는 단어 수준 타이밍을 초래한다. 우리는 모델이 이미 두 스타일을 모두 인코딩하고 있음을 보여주며, 과제는 통제된 활성화에 있음을 밝힌다. 병렬 축어적/의도적 전사 쌍으로 훈련된 커버리지 인식 디코더 작업 토큰을 사용하여, 오직 영어로만 훈련되었음에도 불구하고 독일어 비유창성 F1을 제로샷으로 10%에서 79%로 향상시킨다. 완전한 영어 전용 미세 조정은 두 언어 모두에서 축어적 정확도, 비유창성 탐지, 및 의도 모드 품질에 있어 모든 기준선을 능가한다. 또한, 우리는 강제 정렬 기준선을 넘어 비유창한 발화에 대한 단어 수준 타임스탬프를 개선하는 지도 크로스-어텐션 미세 조정을 추가로 도입한다. 마지막으로, 우리는 고품질의 표준 축어적 전사를 통해 음성 코퍼스의 확장 가능한 생성 및 강화를 가능하게 하는 새로운 작업인 verbatimize를 제안한다.
English
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.