ChatPaper.aiChatPaper

EchoWM:开放且可进入的全模态世界模型

EchoWM: Open and Enterable Omnimodal World Models

August 24, 2026
作者: Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan
cs.AI

摘要

我们提出了 EchoWM,一种用于可进入生成式媒体的全模态世界模型,能够响应连续导航并同时联合生成 720p 视频、环境音、音乐和语音。我们将交互围绕相机意图进行组织:在第一人称场景中,相机意图指定观察者的运动;而在第三人称场景中,相机与角色之间的动态从数据中学习,无需视角专用控制器。离散命令和连续姿态被映射到共享的度量尺度相对六自由度(6-DoF)轨迹,并通过数据集级校准来保持跨异构数据的运动幅度。为了联合学习视听生成与轨迹控制,我们构建了一个互补数据引擎,并采用渐进式训练随后进行自回归后训练的方式,以支持长时程生成。大量评估表明,该模型在公开世界模型基准上实现了强大的轨迹跟随能力和较高的视觉质量,支持跨多种主体的第一人称和第三人称交互,并在长时程生成中保持环境音与语音的同步。
English
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.