ChatPaper.aiChatPaper

潜在変数としての書き起こしポリシー:単語レベルのタイミングで制御可能な逐語ASRの活性化

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

July 21, 2026
著者: Laurin Wagner, Mario Zusag, Bernhard Thallinger
cs.AI

要旨

不均一に注釈付けされたデータで訓練された現代のASRモデルは、転写スタイル(逐語的 vs. 意図的)を制御されていない潜在変数として扱い、その結果、測定可能な復号の不安定性、評価の交絡(報告されたWERの最大60%がスタイルの不一致に起因する)、および信頼性の低い単語レベルのタイミングを引き起こす。我々は、モデルが既に両方のスタイルを符号化していること、課題は制御された活性化にあることを示す。カバレッジを考慮したデコーダタスクトークンを、並列な逐語的/意図的転写ペアで訓練することにより、英語のみの訓練にもかかわらず、ドイツ語の非流暢性F1をゼロショットで10%から79%に向上させる。完全な英語のみの微調整は、両言語において、逐語的正確性、非流暢性検出、および意図モード品質においてすべてのベースラインを上回る。さらに、強制アライメントのベースラインを超えて、非流暢な音声に対する単語レベルのタイムスタンプを改善する教師ありクロスアテンション微調整を導入する。最後に、高品質な正準逐語転写を用いた音声コーパスのスケーラブルな作成と拡充を可能にする新しいタスクであるverbatimize(逐語化)を提案する。
English
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.