SolarWM: 장기 비디오 세계 모델을 위한 공개 데이터 및 확장 가능한 훈련
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
September 2, 2026
저자: Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
cs.AI
초록
우리는 SolarWM을 소개한다. SolarWM은 데이터 준비부터 장기 지평(long-horizon) 추론에 이르기까지 인터랙티브 비디오 월드 모델을 구축하기 위한 완전 개방형 기반이다. 이질적인 데이터 소스와 비디오 백본에 걸친 학습은 어려운 과제이다. 데이터셋마다 시간적 스케일, 카메라 기하 구조, 시각 품질, 모션, 캡셔닝 스타일이 다르며, 비디오 생성 모델은 서로 다른 표현 방식과 아키텍처를 사용한다. 따라서 단순한 데이터 혼합과 모델별 맞춤 구현은 비일관된 지도 신호(supervision)를 초래하고, 결과의 재현과 비교를 어렵게 만든다. SolarWM은 이러한 결합 문제를 재구성 가능한 다중 소스 데이터 엔진과 백본-네이티브 적응 프레임워크로 해결한다. 이 엔진은 10개 데이터셋에서 추출한 143만 개의 표준 캐노니컬 클립을 시각 관측, 메트릭 카메라 기하, 캡션, 품질 메타데이터, 선택 결정, 출처 정보를 포함하는 통합된 프레임 정렬 명세로 변환하며, 소스 처리를 혼합 구성으로부터 분리한다. 공유된 카메라 컨디셔닝, 학습, 추론 인터페이스 하에서 우리는 Wan2.2, LTX-2.5, MiniMax-H3 기반의 5B~33B 규모 네 가지 모델을 각자의 고유한 표현과 목적 함수를 유지한 채 구현한다. 통합된 3단계 레시피는 양방향 적응, 교사 강제(teacher-forced) 자기회귀 초기화, 분포 정합 증류를 결합한다. 그 결과 얻어진 인과적(causal) 모델들은 5초 길이의 시퀀스만으로 학습된 이후에도 수 분에서 수 시간에 이르는 롤아웃 전반에 걸쳐 실시간 상호작용을 가능하게 한다. SolarWM은 생성된 데이터, 파이프라인, 레시피, 가중치, 프레임워크를 모두 공개함으로써 인터랙티브 월드 모델 연구를 위한 재현 가능하고 확장 가능한 기반을 제공한다.
English
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.