ChatPaper.aiChatPaper

Ex-Omni-2D:具原生視覺存在感之表達性全模態對話模型

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

August 11, 2026
作者: Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
cs.AI

摘要

全模態對話模型能理解多模態輸入並生成口語回覆,然而其回覆在視覺上仍缺乏實體。我們提出 Ex-Omni-2D,一個全模態對話框架,可生成包含文字、個人化語音與參考條件化影片的協同回覆。給定多模態查詢、參考影像與參考音訊,此模型先預測描述場景、情感與動作的結構化視覺思維計劃(VTP),隨後生成回覆文字與原生多碼本語音單元。這些語音單元構成共享的聲學-時間介面:它們被解碼為語音並與影片畫格線上對齊。此介面使回覆路徑與虛擬化身路徑得以從異質的語音、對話與化身影片資料中學習,避免了對大規模查詢-文字-語音-影片監督資料的需求。全序列影片生成器作為主要教師模型。為實現高效的增量生成,我們進一步將其蒸餾為少步區塊因果串流學生模型,其前綴串流機制在連續區塊間攜帶乾淨的潛在表示,以減少後期區塊的累積退化。在四步推論下,完整的四 GPU 管線在 400×720/720×400 解析度下達成 1.293 的端到端即時因子(RTF),提供了一個兼具品質與效率的實用操作點。
English
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.