ChatPaper.aiChatPaper

EchoWM: オープンかつ入場可能なオムニモーダル世界モデル

EchoWM: Open and Enterable Omnimodal World Models

August 24, 2026
著者: Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan
cs.AI

要旨

本稿では、連続的なナビゲーションに応答しながら720p映像、環境音、音楽、音声を同時に生成する、入場可能な生成メディアのためのオムニモーダルワールドモデルであるEchoWMを提案する。インタラクションはカメラ意図を中心に構成される。一人称シーンではカメラ意図が観察者の動作を指定し、三人称シーンではカメラとキャラクターのダイナミクスを、視点固有のコントローラなしでデータから学習する。離散コマンドと連続ポーズは、共有のメートルスケール相対6自由度軌跡にマッピングされ、データセットレベルのキャリブレーションにより、異種データ間で移動量が維持される。音声視覚生成と軌跡制御を同時に学習するため、補完的なデータエンジンを構築し、長期生成のためにプログレッシブトレーニングに続く自己回帰的事後トレーニングを採用する。広範な評価により、EchoWMが公開ワールドモデルベンチマークにおいて優れた軌跡追従性能と高い画質を達成し、多様な対象に対する一人称・三人称の両方のインタラクションをサポートし、長期生成にわたって環境音と音声の同期を維持することを示す。
English
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.