ChatPaper.aiChatPaper

Super Star: 디지털 휴먼을 위한 스트리밍 실시간 인터랙티브 에이전트를 향하여

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

July 22, 2026
저자: Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
cs.AI

초록

기존의 동반 음성 제스처 생성 방법들은 주로 오프라인 환경에서 연구되어 왔으며, 이 경우 완전한 음성 구간으로부터 제스처가 합성된다. 그러나 실제 시나리오에서의 대화형 디지털 휴먼은 엄격한 지연 시간 제약 하에서 현재 이용 가능한 응답 오디오만을 사용하여 음성과 동기화된 제스처를 온라인으로 생성해야 한다. 그 결과, 기존 방법들은 미래 음성 정보에 의존하거나 상당한 추론 지연을 초래하기 때문에 실시간 상호작용에 적합하지 않다. 본 논문에서는 대화형 디지털 휴먼을 위한 온라인 동반 음성 제스처 생성을 정식화하고, 스트리밍 음성 응답 모듈과 온라인 제스처 생성 모듈을 결합한 실시간 대화형 프레임워크를 제안한다. 구체적으로, 제스처 생성기는 인과적 다중 모달 자기회귀 모델로 설계되어 스트리밍 응답 음성과 모션 이력으로부터 신체 모션을 예측함으로써, 미래 음성에 대한 접근 없이 저지연 및 음성 정렬 제스처 합성을 가능하게 한다. 이러한 설정을 지원하기 위해, 우리는 가상 동반자 시나리오에 맞춰진 오프라인 데이터 합성 파이프라인을 추가로 제안하는데, 이는 주제 및 감정 인식 대상 말뭉치를 활용하여 다양한 인간-에이전트 대화를 구축한 후 에이전트 응답에 조건화된 동반 음성 제스처를 생성한다. 더 나아가, 오프라인 데이터 구축과 온라인 배포 간의 격차를 해소하기 위해, 온라인 상호작용 중 수집된 사용자 피드백을 데이터 생성 과정에 통합하는 자기 진화 훈련 루프를 구축하여 사용자 선호도에 대한 지속적인 적응을 가능하게 한다. 광범위한 실험을 통해 우리의 프레임워크가 경쟁력 있는 기존 기준선들보다 우수한 지연-품질 트레이드오프, 더 강력한 음성-모션 동기화, 그리고 더 높은 사용자 선호도를 달성함을 입증한다. 프로젝트 페이지: https://super-star-2026.github.io/
English
Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned gesture synthesis without access to future speech. To support this setting, we further propose an offline data synthesis pipeline tailored to virtual companion scenarios, which leverages topic- and emotion-aware subject corpora to construct diverse human-agent dialogues and then generates co-speech gestures conditioned on the agent responses. Moreover, to bridge the gap between offline data construction and online deployment, we establish a self-evolving training loop by incorporating user feedback collected during online interaction into the data generation process, enabling continual adaptation to user preferences. Extensive experiments demonstrate that our framework achieves superior better latency-quality trade-off, stronger speech-motion synchronization, and higher user preference than competitive existing baselines. Project Page: https://super-star-2026.github.io/