ChatPaper.aiChatPaper

說話者分離的區塊迴歸用於音節分詞

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

July 5, 2026
作者: Ryota Komatsu, Kota Kawakita, Takuma Okamoto, Takahiro Shinozaki
cs.AI

摘要

無監督音節分詞旨在從原始語音中學習離散的音節標記,以捕捉與潛在語言內容相關的結構。近年來的音節分詞方法採用預訓練HuBERT的教師-學生蒸餾技術,將潛在語音表徵組織成音節片段。然而,當使用語句層級的交叉熵目標進行訓練時,模型會預測說話者身份而非語言內容,從而損害音節標記的純淨性。為解決此問題,我們提出了一種說話人解耦的音節分詞器,該模型在固定長度區塊內將受到說話者擾動的學生表徵回歸至乾淨的教師目標。實驗結果顯示,我們提出的方法在音節邊界檢測與音節片段聚類中達到了最先進的性能。此外,基於我們的音節標記訓練的語音語言模型,在句法與語義理解上相較於音素層級的SpiRit-LM實現了7%的相對提升。
English
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.