ChatPaper.aiChatPaper

Puppeteer:以物件為基礎的姿態感知共語手勢生成

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

August 31, 2026
作者: Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
cs.AI

摘要

生成兼具時間連貫性、與語音語意對齊,並以周遭物件為基礎的伴隨語音手勢,仍是一項挑戰。先前的語音驅動手勢模型強調音訊—手勢對齊,但未明確考量姿勢限制或周遭物件,因而無法捕捉身體手勢與物理空間之間的固有關聯。我們提出 Puppeteer,一個在因果潛在空間中運作、具姿勢感知且以物件為基礎的伴隨語音手勢擴散模型。我們將長手勢分解為結構化基元,並學習一個因果變分自編碼器,將其編碼為依時間排序的潛在符元,每個潛在符元僅依賴過去。接著,我們直接在因果潛在空間中進行條件式擴散,以語音訊號、動作歷史、初始姿勢參考與物件幾何為條件,合成物理上一致的手勢。這種依時間排序的潛在表述可實現明確的時間控制,並支援手勢補間與手勢補全等任務。為了在現有衡量方式之外更好地評估伴隨語音手勢合成,我們引入針對此任務量身打造的新評估指標。我們也建立了 SceneGes,這是首個經精心策劃、包含具身伴隨語音手勢及對應 3D 物件的合成 3D 資料集,使以物件為基礎的手勢生成得以實現。實驗顯示,Puppeteer 能生成比先前方法更多樣且時間同步的手勢,同時實現以物件為基礎的手勢合成。
English
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.