ChatPaper.aiChatPaper

DuplexGen: 人間とAIのターンテイキング対話の適応的合成

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

July 28, 2026
著者: Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-Tür
cs.AI

要旨

ターンテイキングはフルデュプレックス対話の中核的構成要素である。どのターンテイキング行動が適切かはシナリオによって異なるが、現在のモデルは文脈に関係なく単一の規範を適用する。この限界は訓練データに起因する。人間同士の音声コーパスは自然なタイミング現象を捉えているが、役割のグラウンディングやシナリオ固有の規範はほとんど提供しない。一方、ヒューリスティックまたはプロンプトによる合成手法は、人間の選好に基づかずにターンテイキング行動を注入する。我々は、少数のスロットレベルにおける人間の選好アノテーションに対してLLMの予測を較正することで、シナリオ適応型ターンテイキングを持つ対話を生成するフレームワークDuplexGenを導入する。6つの協力的および競争的タスクにおいて、人間のターンテイキング選好は系統的に異なり、DuplexGenは、較正されていないプロンプティングや一般的な人間同士のデータのみでの訓練よりも、それらの選好と大幅に密接に一致する。また、DuplexGenが生成したデータで訓練されたフルデュプレックスモデルは、特徴的で人間に好まれるターンテイキング行動を示す。これらの結果は、ターンテイキング合成をシナリオ固有にするのは、コーパスの規模やプロンプト設計だけではなく、人間による較正であることを示している。
English
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.