ChatPaper.aiChatPaper

EchoWM:開放且可進入的全模態世界模型

EchoWM: Open and Enterable Omnimodal World Models

August 24, 2026
作者: Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan
cs.AI

摘要

我們提出 EchoWM,這是一個用於可進入式生成媒體的全模態世界模型,能夠回應連續導覽,同時聯合生成 720p 影片、環境音、音樂與語音。我們圍繞相機意圖組織互動:在第一人稱場景中,它指定觀察者的運動;而在第三人稱場景中,相機-角色動態則從資料中學習,無需視角專用控制器。離散指令與連續姿態被對應到共享的公尺度相對六自由度軌跡,並透過資料集層級校正來保留異質資料中的運動幅度。為了聯合學習視聽生成與軌跡控制,我們建構了互補性的資料引擎,並採用漸進式訓練,接著進行自迴歸後訓練以支援長時程生成。廣泛的評估顯示,EchoWM 在公開世界模型基準上實現了強大的軌跡跟隨能力與高視覺品質,支援多種主體的第一人稱與第三人稱互動,並在長時程生成中維持同步的環境音與語音。
English
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.