EchoWM: 개방형이자 진입 가능한 옴니모달 세계 모델
EchoWM: Open and Enterable Omnimodal World Models
August 24, 2026
저자: Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan
cs.AI
초록
본 논문에서는 연속적인 내비게이션에 반응하면서 720p 비디오, 환경음, 음악, 음성을 동시에 생성하는 진입 가능한 생성 미디어(enterable generative media)를 위한 옴니모달 월드 모델인 EchoWM을 제시한다. 우리는 상호작용을 카메라 의도(camera intent)를 중심으로 구성한다: 1인칭 장면에서는 관찰자의 움직임을 지정하고, 3인칭 장면에서는 카메라-캐릭터 역학을 뷰별 컨트롤러 없이 데이터로부터 학습한다. 이산 명령과 연속 자세는 공유된 미터법 스케일의 상대적 6-DoF 궤적으로 매핑되며, 데이터셋 수준의 보정을 통해 이질적 데이터 간에도 움직임의 크기가 보존된다. 오디오-비주얼 생성과 궤적 제어를 공동으로 학습하기 위해 상보적 데이터 엔진을 구축하고, 점진적 훈련(progressive training) 이후 자기회귀적 사후 훈련(autoregressive post-training)을 채택하여 장기간 생성(long-horizon generation)을 수행한다. 광범위한 평가를 통해 본 모델이 공개 월드 모델 벤치마크에서 강력한 궤적 추종 성능과 높은 시각적 품질을 달성하며, 다양한 피사체에 걸쳐 1인칭 및 3인칭 상호작용을 모두 지원하고, 장기간 생성 과정에서 동기화된 환경음과 음성을 유지함을 입증한다.
English
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.