ChatPaper.aiChatPaper

Ex-Omni-2D: ネイティブな視覚プレゼンスを備えた表現力豊かなオムニモーダル対話モデル

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

August 11, 2026
著者: Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
cs.AI

要旨

オムニモーダル対話モデルは、マルチモーダル入力を理解し、音声応答を合成できるが、その応答は視覚的には身体性を欠いたままである。本稿では、テキスト、パーソナライズされた音声、および参照条件付きビデオからなる協調応答を生成するオムニモーダル対話フレームワークEx-Omni-2Dを提案する。マルチモーダルクエリ、参照画像、参照オーディオが与えられると、本モデルはシーン、感情、動作を記述する構造化された視覚思考プラン(VTP)を予測し、続いて応答テキストとネイティブなマルチコードブック音声ユニットを予測する。これらのユニットは共有音響・時間インターフェースを構成し、音声にデコードされると同時にビデオフレームとオンラインで整列される。このインターフェースにより、応答経路とアバター経路を異種の音声データ、対話データ、アバタービデオデータから学習でき、大規模なクエリ‐テキスト‐音声‐ビデオの教師データを必要としない。フルシーケンスのビデオジェネレータが主要な教師モデルとして機能する。効率的な増分生成のために、これを数ステップのブロック因果的ストリーミング生徒モデルへさらに蒸留する。そのプレフィックスストリーミング機構は、連続するチャンク間でクリーンな潜在表現を保持し、後期チャンクにおける累積的な劣化を低減する。4ステップ推論により、4GPU構成の完全なパイプラインは、400×720/720×400の解像度でエンドツーエンドのRTF 1.293を達成し、実用的な品質と効率の動作点を提供する。
English
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.