話者分離チャンク単位回帰による音節トークン化
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
July 5, 2026
著者: Ryota Komatsu, Kota Kawakita, Takuma Okamoto, Takahiro Shinozaki
cs.AI
要旨
教師なし音節トークン化は、生の音声から潜在的な言語内容に関連する構造を捉える離散的な音節トークンを学習することを目的としている。近年の音節トークン化手法では、事前学習済みHuBERTの教師-生徒蒸留を利用し、潜在的な音声フレーム表現を音節セグメントに整理する。しかし、発話レベルの交差エントロピー目的関数で学習すると、モデルは言語内容ではなく話者識別を予測するため、音節トークンの純度が損なわれる。この問題に対処するため、固定長チャンク内で話者に摂動を加えた生徒表現をクリーンな教師ターゲットに回帰する、話者非絡み合い音節トークン化器を提案する。実験結果は、提案手法が音節境界検出と音節セグメントクラスタリングにおいて最先端の性能を達成することを示している。さらに、提案する音節トークンで学習した音声言語モデルは、音素レベルのSpiRit-LMと比較して、統語的・意味的理解において7%の相対的な改善を達成する。
English
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.