音韻活性化マッピングによる音素セグメンテーションと認識
Phone Segmentation and Recognition through Phonological Activation Mapping
July 10, 2026
著者: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
cs.AI
要旨
音素セグメンテーションと認識は本質的に関連するタスクであるにもかかわらず、現代の手法では通常これらを別々にモデル化している。我々は、音韻構造が自己教師あり音声モデル(S3Ms)の表現にすでに潜在しており、両タスクを解決するためにはモデルを誘導するだけで十分であると主張する。我々はS3Mベースの音韻活性化マッピング(SPAM)を活用する。これは各S3M表現フレームを、有声性や鼻音性などの音韻特徴活性化のベクトルにマッピングするものである。SPAMの上に、我々は2つのシンプルで効果的かつ軽量な勾配降下不要の予測ヘッド(認識ヘッドとセグメンテーションヘッド)を導入する。本手法は1分未満の音声表記を必要とし、訓練時に未学習の音素に対しても汎化する。多様なデータセットにわたって、本アプローチは強力なセグメンテーションおよび認識性能を達成する。
English
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.