ChatPaper.aiChatPaper

AlayaWorld: 장기적이고 플레이 가능한 비디오 월드 생성

AlayaWorld: Long-Horizon and Playable Video World Generation

July 7, 2026
저자: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
cs.AI

초록

게임 세계는 전통적으로 노동 집약적인 생산 파이프라인을 통해 구축되어 왔으며, 이로 인해 개발 비용이 많이 들고, 맞춤화가 어렵고, 배포 후 수정에 많은 비용이 소요된다. 최근 비디오 월드 모델의 발전은 근본적으로 다른 패러다임을 제공한다. 이러한 모델들은 가상 환경의 모든 구성 요소를 명시적으로 제작하는 대신, 현재 세계 상태와 사용자 상호작용에 기반하여 미래 관찰을 자기회귀적으로 합성함으로써, 플레이 가능한 세계를 온라인으로 생성할 수 있게 한다. 게임플레이 기록과 실제 세계 비디오 모두에 대해 훈련된 이 모델들은 다양한 시각적 외관과 물리적 동역학을 포착할 수 있으며, 이는 게임을 넘어 구현된 지능을 포함하는 대화형 애플리케이션에 새로운 기회를 열어준다. 본 논문에서는 대화형 생성 세계를 구축하기 위한 풀스택 오픈소스 프레임워크인 AlayaWorld를 제시한다. AlayaWorld는 개방형 실시간 상호작용을 가능하게 하여, 사용자가 자유롭게 탐색하고 전투, 주문 시전, 몬스터 소환과 같은 다양한 행동을 수행할 수 있게 한다. 이 프레임워크는 데이터 준비, 모델 아키텍처, 모델 훈련, 추론 가속화 및 배포에 이르는 완전한 개발 과정을 모듈식 및 확장 가능한 아키텍처 내에서 통합한다. 프레임워크와 함께, 우리는 재현 가능한 파이프라인, 참조 구현, 평가 도구 및 포괄적인 문서를 공개함으로써, 생성적 월드 모델의 향후 연구 및 실시간 애플리케이션을 위한 실용적인 기반을 마련한다.
English
Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than explicitly authoring every component of a virtual environment, these models autoregressively synthesize future observations conditioned on the current world state and user interactions, enabling playable worlds to be generated online. Trained on both gameplay recordings and real-world videos, they can capture diverse visual appearances and physical dynamics, opening new opportunities for interactive applications beyond gaming, including embodied intelligence. In this paper, we present AlayaWorld, a full-stack open-source framework for building interactive generative worlds. AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning. The framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-within a modular and extensible architecture. Alongside the framework, we release reproducible pipelines, reference implementations, evaluation tools, and comprehensive documentation, establishing a practical foundation for future research and real-time applications of generative world models.