Puppeteer: 物体に基づく姿勢考慮型共発話ジェスチャ生成
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
August 31, 2026
著者: Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
cs.AI
要旨
時間的に一貫し、発話と意味的に整合し、周囲の物体に接地した共発話ジェスチャーの生成は依然として困難である。従来の音声駆動型ジェスチャーモデルは音声とジェスチャーの整合性を重視するが、姿勢制約や周囲物体を明示的に考慮しておらず、身体ジェスチャーと物理空間の本質的な相関を捉えられていない。本稿では、因果潜在空間で動作する、姿勢を考慮し物体に接地した共発話ジェスチャー拡散モデルPuppeteerを提案する。我々は長いジェスチャーを構造化プリミティブに分解し、それらを時間順に並んだ潜在トークンへ符号化する因果変分オートエンコーダを学習する。各潜在トークンは過去のみに依存する。次に、因果潜在空間において直接条件付き拡散を実行し、音声信号、運動履歴、初期姿勢参照、物体ジオメトリを条件として、物理的に整合したジェスチャーを合成する。この時間順潜在表現により、明示的な時間制御が可能になり、ジェスチャー補間やジェスチャー補完などのタスクを支援する。既存の評価尺度を超えて共発話ジェスチャー合成をより適切に評価するため、本タスクに特化した新しい評価指標を導入する。また、物体接地型ジェスチャー生成を可能にする、身体化された共発話ジェスチャーと対応する3D物体からなる最初のキュレーション済み合成3DデータセットSceneGesを作成した。実験により、Puppeteerは従来手法よりも多様で時間的に同期したジェスチャーを生成し、かつ物体接地型ジェスチャー合成を可能にすることが示された。
English
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.