ChatPaper.aiChatPaper

Puppeteer: 객체에 근거한 자세 인지 동시 발화 제스처 생성

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

August 31, 2026
저자: Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
cs.AI

초록

시간적으로 일관되고 발화와 의미적으로 정렬되며 주변 객체에 기반을 둔 발화 동반 제스처를 생성하는 것은 여전히 어려운 과제이다. 기존의 발화 기반 제스처 모델들은 오디오-제스처 정렬을 강조하지만 자세 제약이나 주변 객체를 명시적으로 고려하지 않아, 신체 제스처와 물리적 공간 사이의 내재적 상관관계를 포착하지 못한다. 우리는 인과적 잠재 공간에서 동작하는 자세 인식 및 객체 기반 발화 동반 제스처 확산 모델인 Puppeteer를 제안한다. 우리는 긴 제스처를 구조화된 프리미티브로 분해하고, 이를 시간 순서가 있는 잠재 토큰으로 인코딩하는 인과적 변분 오토인코더를 학습하며, 각 토큰은 과거에만 의존한다. 그다음, 음성 신호, 동작 이력, 초기 자세 참조, 객체 형상을 조건으로 하여 인과적 잠재 공간에서 직접 조건부 확산을 수행함으로써 물리적으로 일관된 제스처를 합성한다. 이러한 시간 순서 잠재 표현은 명시적 시간 제어를 가능하게 하고 제스처 중간 보간 및 제스처 완성과 같은 작업을 지원한다. 기존 척도를 넘어 발화 동반 제스처 합성을 더 잘 평가하기 위해, 우리는 이 작업에 맞춘 새로운 평가 지표를 도입한다. 또한 객체 기반 제스처 생성을 가능하게 하는, 체화된 발화 동반 제스처와 대응하는 3D 객체로 구성된 최초의 정제된 합성 3D 데이터셋인 SceneGes를 구축한다. 실험은 Puppeteer가 기존 방법보다 더 다양하고 시간적으로 동기화된 제스처를 생성하면서 객체 기반 제스처 합성을 가능하게 함을 보여준다.
English
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.