ABot-World-0: 単一のデスクトップGPU上での無限インタラクティブワールドのロールアウト
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
July 21, 2026
著者: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
cs.AI
要旨
我々はABot-World-0を提案する。これは、リアルタイムかつ長期的な閉ループインタラクションのための行動条件付きビデオワールドモデルであり、AAAゲーム、シミュレーションエンジン、インターネット動画にわたるマルチソースデータインフラに支えられ、制御可能な世界ダイナミクスを学習する。WorldExplorerは、訓練フィードバックに導かれたエージェント主導の収集を実行し、統一パイプラインは14の決定論的品質チェック、VLMベースの評価、同期された行動・テキスト注釈を適用する。我々は、教師強制とODE蒸留を通じて双方向行動条件付き教師モデルを因果的生徒モデルへと段階的に蒸留し、LongForcingを導入することで、長期の生徒自己ロールアウトを拡張時間軸の教師モデルと整合させ、累積的な分布シフトと自己回帰ドリフトを軽減する。生のキーボード行動は、シーン探索や三人称キャラクターインタラクションのための統一制御インターフェースを提供し、参照キャラクターメモリは三人称ロールアウト中にアイデンティティ一貫性のための持続的な外見的手がかりを提供する。デプロイメントのために、我々は軽量VAEデコーダ、効率的なアテンション、メモリ認識スケジューリング、低ビットDiT推論を備えたストリーミング推論スタックを共同設計する。最適化された低ビット構成において、ABot-World-0は単一のNVIDIA RTX 5090デスクトップGPU上で最大16FPSの720P動画をストリーミングし、アクションから初フレームまでのレイテンシは1.2秒、ピークVRAMは約19GiBである。WorldRoamBenchおよび拡張インタラクティブロールアウトにおける実験は、競争力のある制御可能性と、首尾一貫した長期的な世界の進化を示している。
English
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.