ChatPaper.aiChatPaper

轉錄策略作為潛在變數:利用詞級時間啟動可控的逐字語音辨識

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

July 21, 2026
作者: Laurin Wagner, Mario Zusag, Bernhard Thallinger
cs.AI

摘要

現代自動語音辨識模型在異質標註資料上訓練時,會將轉寫風格(逐字對比意圖)視為未受控制的潛在變數,導致可測量的解碼不穩定性、評估混淆(多達60%的詞錯誤率差異源自風格不匹配),以及不可靠的詞層級時間戳記。我們證明模型實際上已編碼兩種風格;問題在於如何受控地啟用。透過在平行逐字/意圖轉寫配對上訓練的覆蓋感知解碼器任務標記,我們在僅接受英文訓練的情況下,將德語不流暢F1分數從10%零樣本提升至79%。純英文微調在逐字準確率、不流暢偵測及兩種語言的意圖模式品質上,均超越所有基準模型。我們進一步引入監督式交叉注意力微調,此法改善不流暢語音上的詞層級時間戳記,表現優於強制對齊基準。最後,我們提出「逐字化」這項新任務,可擴展地建立及豐富語料庫,並產出高品質的標準逐字轉寫。
English
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.