SolarWM:面向長時域視頻世界模型的開放數據與可擴展訓練
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
September 2, 2026
作者: Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
cs.AI
摘要
我們推出 SolarWM——一個完全開放的基礎架構,用於建構從資料準備到長時程推論的互動式影片世界模型。在異質資料來源與各種影片骨幹網路上訓練極具挑戰性:不同資料集的時間尺度、相機幾何、視覺品質、動態與描述文字風格各不相同,而影片生成器也各自採用不同的表徵與架構。因此,單純混合資料或針對特定模型作個別實作,會產生不一致的監督訊號,使結果難以重現與比較。為解決此耦合問題,SolarWM 結合了可重構的多來源資料引擎,以及以骨幹網路原生特性為本位的適配框架。該引擎將來自 10 個資料集的 143 萬個標準化片段,轉換為統一且逐幀對齊的資料契約,涵蓋視覺觀測、度量相機幾何、描述文字、品質後設資料、選擇決策與來源出處,同時將來源處理與混合資料建構解耦。在共享的相機條件化、訓練與推論介面下,我們基於 Wan2.2、LTX-2.5 與 MiniMax-H3 實例化了四個 5B 至 33B 參數規模的模型,同時保留其原生表徵與訓練目標。一套統一的三階段訓練配方結合了雙向適配、教師強迫自迴歸初始化與分佈匹配蒸餾。所得的因果模型僅以 5 秒序列訓練後,即可在長達數分鐘至數小時的推演中支援即時互動。透過公開釋出上述資料、流程、配方、權重與框架,SolarWM 為互動式世界模型研究提供了可重現且可擴展的基礎。
English
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.