ChatPaper.aiChatPaper

Foneemsegmentatie en -herkenning via fonologische activeringsmapping

Phone Segmentation and Recognition through Phonological Activation Mapping

July 10, 2026
Auteurs: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
cs.AI

Samenvatting

Foneemsegmentatie en -herkenning zijn inherent gerelateerde taken, maar moderne benaderingen modelleren ze doorgaans apart. Wij stellen dat fonetische structuur al latent aanwezig is in de representaties van zelfgesuperviseerde spraakmodellen (S3Ms), en dat men deze slechts hoeft te sturen om beide taken op te lossen. We maken gebruik van op S3M gebaseerde fonologische activeringsmapping (SPAM), die elk S3M-representatieframe toewijst aan een vector van fonologische kenmerkactivaties, zoals stemhebbendheid en nasaliteit. Bovenop SPAM introduceren we twee eenvoudige maar effectieve, lichtgewicht, gradiëntafdalingsvrije voorspellingskoppen: een herkenningskop en een segmentatiekop. Onze methode vereist minder dan een minuut aan fonetische transcripties en generaliseert naar ongeziene fonemen tijdens de training. Over een diverse reeks datasets behaalt onze aanpak sterke segmentatie- en herkenningsprestaties.
English
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.