Wonder: 더 나은 비디오 월드 모델
Wonder: Video World Model Done Better
July 28, 2026
저자: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
cs.AI
초록
우리는 실시간 카메라 제어가 가능한 세계 탐험을 위한 범용 비디오 월드 모델인 Wonder를 제시한다. Wonder는 이미지 또는 조건부 비디오가 주어지면 사용자가 카메라를 움직여 대화형으로 탐색하고, 보지 못한 영역을 발견하며, 이전에 관찰된 영역을 실시간으로 그리고 장기적 시간 범위에 걸쳐 재방문할 수 있는 플레이 가능한 세계를 구축한다. 이러한 능력을 달성하려면 제어 방법, 메모리 메커니즘 및 훈련 전략의 시스템 수준 공동 설계가 필요하다. 우리는 렌더링을 통해 공간적으로 정렬된 움직임과 방향 신호를 제공하는 밀집 좌표 필드를 이용한 새로운 카메라 조건화 방식을 도입하여, 모델이 카메라 움직임을 시각적 증거로 직접 해석할 수 있도록 한다. 증가하는 생성 컨텍스트에 걸쳐 빠르고 정확한 메모리 검색을 지원하기 위해, 우리는 효율적인 희소 어텐션 기반 메모리 메커니즘을 제안한다. 이를 통해 모델은 실제 컨텍스트 길이와 관계없이 추론 시점에 소수의 관련 컨텍스트 토큰 집합에 선택적으로 주목할 수 있다. 또한, 우리는 자기 강제 방식 증류 파이프라인을 개선하기 위한 여러 기술을 개발하여, 학생 모델이 제어 신호를 존중하는 능력을 향상시키고, 교사 모델의 다양한 생성 모드와 장기 기억을 유지한다. 이러한 구성 요소들은 함께 작동하여 Wonder가 긴 롤아웃 동안 일관된 기하학, 외관 및 역학을 유지하면서 16FPS로 다양하고 분 단위의 비디오를 합성할 수 있게 한다. 이미지-비디오 생성 외에도, Wonder는 자연스럽게 비디오 조건부 생성을 지원하여 기존의 동적 장면을 실시간으로 재촬영할 수 있게 한다.
English
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.