ChatPaper.aiChatPaper

基于音韵激活映射的音素分割与识别

Phone Segmentation and Recognition through Phonological Activation Mapping

July 10, 2026
作者: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
cs.AI

摘要

语音分割与识别本质上是相互关联的任务,但现代方法通常分别对二者进行建模。我们认为,音系结构已隐含在自监督语音模型(S3M)的表征中,只需引导这些模型即可解决这两项任务。我们提出基于S3M的音系激活映射(SPAM)方法,该方法将每个S3M表征帧映射为一个音系特征激活向量,例如发声性、鼻音性等特征。在SPAM基础上,我们引入了两个简单却有效的轻量级无梯度下降预测头:识别头与分割头。该方法仅需不到一分钟的音标转录数据,且能泛化到训练中未见的音素。在多种数据集上,我们的方法均取得了优异的分割与识别性能。
English
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.