ChatPaper.aiChatPaper

透過音韻激活映射之語音分割與辨識

Phone Segmentation and Recognition through Phonological Activation Mapping

July 10, 2026
作者: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
cs.AI

摘要

音素切割與辨識本質上是相互關聯的任務,然而現代方法通常各自獨立建模。我們主張,自監督語音模型(S3Ms)的表示中早已潛藏著音韻結構,僅需引導這些模型即可同時解決兩項任務。我們利用基於S3M的音韻激活映射(SPAM),將每個S3M表示幀映射為一組音韻特徵激活向量,例如清濁音與鼻音化。在此基礎上,我們引入兩種簡單而有效的輕量級、免梯度下降預測頭:辨識頭與切割頭。本方法僅需不到一分鐘的音標轉寫資料,且能泛化至訓練中未見過的音素。在涵蓋多樣化資料集的情況下,我們的方法達到了優異的切割與辨識表現。
English
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.