ChatPaper.aiChatPaper

음운 활성화 매핑을 통한 음소 분할 및 인식

Phone Segmentation and Recognition through Phonological Activation Mapping

July 10, 2026
저자: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
cs.AI

초록

음소 분할과 인식은 본질적으로 관련된 작업이지만, 현대적 접근법은 일반적으로 이들을 별도로 모델링한다. 본 논문은 음성 구조가 이미 자기 지도 음성 모델(S3M)의 표현에 잠재되어 있으며, 두 작업을 모두 해결하기 위해 이를 조정하기만 하면 된다고 주장한다. 우리는 S3M 기반 음운 활성 매핑(SPAM)을 활용하며, 이는 각 S3M 표현 프레임을 유성음 및 비음성과 같은 음운 특징 활성화 벡터로 매핑한다. SPAM 위에 우리는 인식 헤드와 분할 헤드라는 두 가지 단순하지만 효과적인 경량의 경사 하강법 불필요 예측 헤드를 도입한다. 본 방법은 1분 미만의 음성 전사만을 필요로 하며, 훈련 중 보지 못한 음소에도 일반화된다. 다양한 데이터셋에서 우리의 접근법은 강력한 분할 및 인식 성능을 달성한다.
English
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.