ChatPaper.aiChatPaper

Ex-Omni-2D: 고유 시각적 존재감을 갖춘 표현적 옴니모달 대화 모델

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

August 11, 2026
저자: Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
cs.AI

초록

옴니모달 대화 모델은 멀티모달 입력을 이해하고 음성 응답을 합성할 수 있지만, 그 응답은 시각적으로 구현되지 않은 상태로 남아 있다. 본 논문에서는 텍스트, 개인화된 음성, 참조 조건부 비디오로 구성된 조정된 응답을 생성하는 옴니모달 대화 프레임워크인 Ex-Omni-2D를 소개한다. 멀티모달 질의, 참조 이미지, 참조 오디오가 주어지면 모델은 장면, 감정, 동작을 설명하는 구조화된 시각적 사고 계획(VTP)을 예측하고, 이어서 응답 텍스트와 네이티브 다중 코드북 음성 유닛을 생성한다. 이러한 유닛은 공유 음향-시간적 인터페이스를 형성하여 음성으로 디코딩되고 비디오 프레임과 온라인으로 정렬된다. 이 인터페이스를 통해 응답 및 아바타 경로를 이종 음성, 대화, 아바타 비디오 데이터로부터 학습할 수 있어 대규모 질의-텍스트-음성-비디오 지도 학습 데이터가 필요하지 않다. 전체 시퀀스 비디오 생성기가 주요 교사(Teacher) 역할을 한다. 효율적인 증분 생성을 위해 이를 소수의 스텝으로 구성된 블록 인과적 스트리밍 학생(Student)으로 추가 증류하며, 해당 학생의 프리픽스 스트리밍 메커니즘은 연속된 청크 간에 깨끗한 잠재 표현을 전달하여 누적 후반 청크 성능 저하를 줄인다. 4스텝 추론을 통해 전체 4-GPU 파이프라인은 400×720/720×400 해상도에서 엔드투엔드 RTF 1.293을 달성하여 실용적인 품질-효율성 운영 지점을 제공한다.
English
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.