DuplexGen:自適應合成人機輪替對話
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
July 28, 2026
作者: Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-Tür
cs.AI
摘要
話輪轉換(turn-taking)是全雙工互動的核心組成部分。何種話輪轉換行為為適當,會隨情境而異,然而現有模型無論情境為何,皆套用單一規範。此限制源於其訓練資料:人與人對話語料庫捕捉了自然的時序現象,但缺乏角色定錨或情境特定規範;而啟發式或提示驅動的合成方法雖能注入話輪轉換行為,卻未以人類偏好為基礎。我們提出 DuplexGen,一個透過將大型語言模型(LLM)預測對齊於少量槽位層級人類偏好標註,來生成具情境適應性話輪轉換對話的框架。在六項合作與競爭任務中,人類的話輪轉換偏好呈現系統性差異,而 DuplexGen 相較於未經校準的提示方法或僅以通用人與人資料訓練之模型,能顯著更貼近這些偏好;以 DuplexGen 生成資料訓練的全雙工模型,展現出獨特且為人類所偏好的話輪轉換行為。這些結果顯示,人類校準(而非語料庫規模或提示設計本身)才是使話輪轉換合成得以具備情境特定性的關鍵。
English
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.