AlayaWorld: 대화형 장기 지평 세계 모델링 — 전체 기술 보고서
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
July 20, 2026
저자: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
cs.AI
초록
기존의 비디오 게임 개발이 자산 제작, 애니메이션, 물리 및 프로그래밍을 위한 노동 집약적 파이프라인에 의존하는 것과 달리, 비디오 월드 모델은 사용자 입력으로부터 즉시 상호작용 가능한 환경을 생성한다. 이를 통해 텍스트, 이미지 또는 비디오로부터 맞춤형이고 탐험 가능하며 지속적으로 진화하는 가상 세계를 만들 수 있다. 이러한 비전을 실현하려면 상호작용, 지속적인 시공간 일관성, 안정적인 장기 생성, 그리고 효율적인 응답이라는 네 가지 긴밀하게 결합된 기능이 필요하다. 우리는 540p 및 720p 해상도에서 초당 24프레임의 비디오를 생성하는 대화형 장기 비디오 월드 모델인 AlayaWorld를 제시한다. 15B 파라미터의 비디오 확산 트랜스포머를 기반으로 구축된 AlayaWorld는 카메라 궤적과 전환 가능한 텍스트 프롬프트 하에서 짧은 잠재 청크를 자기회귀적으로 생성한다. 그 경계 시각적 컨텍스트는 지속적인 싱크 프레임, 압축된 시간적 히스토리, 기하학 정렬 공간 메모리, 그리고 최근 프레임 조건화를 결합한다. 장기적인 드리프트를 줄이기 위해 모델은 자체 롤아웃에서 수집된 손상된 히스토리와 예측 잔차로 훈련된다. 또한 분포 정합 증류, self-forcing++, 및 일관성 증류를 결합한 이산적 자기회귀 증류 공식을 도입하여, 청크당 약 30개의 샘플링 단계를 4단계로 줄인다. iWorld-Bench에서 AlayaWorld는 장기 생성 부문에서 최고 성능을 달성한다. 전체 스택, 오픈소스, 장기 프로젝트로 구상된 AlayaWorld는 향후 대화형 비디오 월드 모델 연구를 위한 확장 가능한 기반을 제공하기 위한 것이다.
English
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.