화자 분리 청크 단위 회귀를 통한 음절 토큰화
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
July 5, 2026
저자: Ryota Komatsu, Kota Kawakita, Takuma Okamoto, Takahiro Shinozaki
cs.AI
초록
비지도 음절 토큰화는 원시 음성으로부터 언어적 내용 관련 잠재 구조를 포착하는 이산 음절 토큰을 학습하는 것을 목표로 한다. 최근 음절 토큰화 방법은 사전 학습된 HuBERT의 교사-학생 증류를 활용하여 잠재 음성 프레임 표현을 음절 단위로 구성한다. 그러나 발화 수준의 교차 엔트로피 목적 함수로 훈련할 경우, 모델은 언어적 내용보다 화자 정체성을 예측하게 되어 음절 토큰의 순수성을 저해한다. 이 문제를 해결하기 위해, 우리는 고정 길이 청크 내에서 화자 교란된 학생 표현을 깨끗한 교사 목표로 회귀시키는 화자 분리 음절 토크나이저를 제안한다. 실험 결과, 제안된 방법이 음절 경계 검출 및 음절 세그먼트 클러스터링에서 최첨단 성능을 달성함을 보여준다. 또한, 우리의 음절 토큰으로 훈련된 음성 언어 모델은 음소 수준의 SpiRit-LM 대비 구문 및 의미 이해에서 7%의 상대적 개선을 달성한다.
English
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.