AlayaWorld:長期的視野を持つ対話可能なビデオワールド生成
AlayaWorld: Long-Horizon and Playable Video World Generation
July 7, 2026
著者: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
cs.AI
要旨
従来、ゲームワールドは労働集約的な制作パイプラインによって構築されてきたため、開発コストが高く、カスタマイズが難しく、デプロイ後の修正にも多大な費用がかかっていた。しかし、近年のビデオ世界モデルの進展により、根本的に異なるパラダイムが登場している。これらのモデルは、仮想環境のあらゆる構成要素を明示的に作り込むのではなく、現在の世界状態とユーザーの操作に基づいて将来の観測を自己回帰的に合成することで、プレイ可能な世界をオンラインで生成することを可能にする。ゲームプレイの記録と実世界の動画の両方で学習されたこのモデルは、多様な視覚的外観と物理的ダイナミクスを捉えることができ、ゲームを超えた応用、すなわち具現化知能を含むインタラクティブなアプリケーションに新たな可能性を拓く。本稿では、インタラクティブな生成的世界を構築するためのフルスタックのオープンソースフレームワークであるAlayaWorldを紹介する。AlayaWorldは、オープンエンドなリアルタイムインタラクションを実現し、ユーザーは自由に移動しながら、戦闘、呪文詠唱、モンスター召喚など多様なアクションを実行できる。本フレームワークは、データ準備、モデルアーキテクチャ、モデル学習、推論高速化、デプロイメントに至るまでの開発プロセス全体を、モジュール化され拡張可能なアーキテクチャのもとに統合している。フレームワークと併せて、再現可能なパイプライン、リファレンス実装、評価ツール、包括的なドキュメントも公開し、生成的世界モデルに関する今後の研究およびリアルタイムアプリケーションのための実践的な基盤を提供する。
English
Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than explicitly authoring every component of a virtual environment, these models autoregressively synthesize future observations conditioned on the current world state and user interactions, enabling playable worlds to be generated online. Trained on both gameplay recordings and real-world videos, they can capture diverse visual appearances and physical dynamics, opening new opportunities for interactive applications beyond gaming, including embodied intelligence. In this paper, we present AlayaWorld, a full-stack open-source framework for building interactive generative worlds. AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning. The framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-within a modular and extensible architecture. Alongside the framework, we release reproducible pipelines, reference implementations, evaluation tools, and comprehensive documentation, establishing a practical foundation for future research and real-time applications of generative world models.