SolarWM: 長期ビデオワールドモデルのためのオープンデータとスケーラブルなトレーニング

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

September 2, 2026
著者: Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
cs.AI

要旨

本稿では、データ準備から長時間にわたる推論までをカバーする、対話型映像世界モデル構築のための完全公開基盤SolarWMを紹介する。異種のデータソースと映像バックボーンにわたる学習は困難である。データセットは時間スケール、カメラジオメトリ、画質、動き、キャプションスタイルが異なり、映像生成モデルもそれぞれ異なる表現とアーキテクチャを採用している。したがって、単純なデータ混合とモデル固有の実装は不整合な教師信号を生み出し、結果の再現と比較を困難にする。SolarWMは、再構成可能なマルチソースデータエンジンとバックボーン固有の適応フレームワークによって、この結合問題に対処する。このエンジンは、10のデータセットに由来する143万本の標準クリップを、視覚観測、メトリックカメラジオメトリ、キャプション、品質メタデータ、選択判断、来歴を網羅した、フレーム整合のとれた統一仕様に変換し、ソース処理と混合データ構築を分離する。共有のカメラ条件付け・学習・推論インターフェースの下で、我々はWan2.2、LTX-2.5、MiniMax-H3に基づく5B〜33Bのモデル4つを、それぞれのバックボーン本来の表現と学習目的を保ったまま実装する。統一された3段階のレシピは、双方向適応、ティーチャーフォーシングによる自己回帰初期化、分布マッチング蒸留を組み合わせる。その結果得られる因果モデルは、学習に用いたのがわずか5秒間のシーケンスのみであるにもかかわらず、数分から数時間に及ぶロールアウトにわたってリアルタイムの対話を可能にする。得られたデータ、パイプライン、レシピ、重み、フレームワークを公開することで、SolarWMは対話型世界モデル研究のための再現可能かつ拡張可能な基盤を提供する。
English
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.
PDF1331September 4, 2026