ChatPaper.aiChatPaper

AlayaWorld: インタラクティブ長期ホライゾン世界モデリング — 完全技術報告書

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

July 20, 2026
著者: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
cs.AI

要旨

従来のゲーム開発がアセット制作、アニメーション、物理演算、プログラミングといった労働集約的なパイプラインに依存していたのとは異なり、ビデオワールドモデルはユーザー入力から即座にインタラクティブな環境を生成する。これにより、テキスト、画像、あるいは動画から、カスタマイズ可能で探索可能、かつ継続的に進化する仮想世界を創り出すことができる。このビジョンを実現するには、インタラクション、持続的な時空間一貫性、安定した長期生成、そして効率的な応答という、密接に関連した4つの能力が必要となる。本稿では、540pおよび720pの解像度で24fpsの動画を生成する、インタラクティブな長期ビデオワールドモデルであるAlayaWorldを提案する。AlayaWorldは、150億パラメータのビデオ拡散トランスフォーマー上に構築されており、カメラ軌道と切り替え可能なテキストプロンプトの下で、短い潜在チャンクを自己回帰的に生成する。その有界視覚コンテキストは、永続的なシンクフレーム、圧縮された時間履歴、幾何学的に整合された空間記憶、および直近フレーム条件付けを組み合わせたものである。長期ドリフトを低減するため、モデルは自身のロールアウトから収集された破損履歴と予測残差を用いて訓練される。さらに、分布一致蒸留、自己強制++、および一致性蒸留を組み合わせた離散自己回帰蒸留定式化を導入し、1チャンクあたりの推論を約30サンプリングステップから4ステップに削減する。iWorld-Benchにおいて、AlayaWorldは長期生成において最高の性能を達成した。フルスタックでオープンソース、かつ長期的なプロジェクトとして構想されたAlayaWorldは、インタラクティブなビデオワールドモデルに関する将来の研究のための拡張可能な基盤を提供することを目的としている。
English
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.