Ex-Omni-2D:具备原生视觉呈现能力的高表现力全模态对话模型
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
August 11, 2026
作者: Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
cs.AI
摘要
全模态对话模型能够理解多模态输入并生成语音回复,但它们的回复在视觉上仍然是无实体的。我们提出Ex-Omni-2D,一种全模态对话框架,能够生成包含文本、个性化语音和参考条件视频的协调回复。给定多模态查询、参考图像和参考音频,模型预测一个描述场景、情感和动作的结构化视觉思维计划(VTP),随后生成回复文本和原生多码本语音单元。这些单元构成一个共享的声学-时间接口:它们被解码为语音并与视频帧在线对齐。该接口使得回复路径和虚拟形象路径能够从异构的语音、对话和虚拟形象视频数据中学习,从而避免了对大规模查询—文本—语音—视频监督的需求。全序列视频生成器充当主要教师。为实现高效的增量生成,我们进一步将其蒸馏为少步块因果流式学生模型,其前缀流式机制在连续块之间携带干净的潜在表示,以减少后续块的累积退化。通过四步推理,完整的四GPU流水线在400×720/720×400分辨率下实现了1.293的端到端RTF,提供了一个实用的质量—效率工作点。
English
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.