ChatPaper.aiChatPaper

阿賴耶世界:互動式長期視野世界建模 ── 完整技術報告

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

July 20, 2026
作者: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
cs.AI

摘要

與傳統依賴勞動密集型管線進行資產製作、動畫、物理及程式開發的電玩遊戲不同,影片世界模型能從使用者輸入即時生成互動環境。它讓我們能透過文字、圖片或影片,創造出可自訂、可探索且持續演進的虛擬世界。實現此願景需要四項緊密耦合的能力:互動性、持久的時空一致性、穩定的長時程生成,以及高效的反應速度。我們提出 AlayaWorld,這是一個互動式長時程影片世界模型,能以 540p 和 720p 的解析度生成每秒 24 幀的影片。該模型建構於 15B 參數的影片擴散變壓器之上,能根據攝影機軌跡及可切換的文字提示,以自迴歸方式生成短暫的潛在區塊。其有界視覺脈絡結合了持久的容器幀、壓縮的時間歷程、符合幾何結構的空間記憶,以及近期幀條件設定。為減少長期漂移,模型會利用其自身推論過程中收集的損毀歷程與預測殘差進行訓練。我們進一步引入一種離散自迴歸蒸餾公式,結合分佈匹配蒸餾、自我強迫++與一致性蒸餾,將每個區塊的推理從約 30 個取樣步驟減少至 4 個步驟。在 iWorld-Bench 上,AlayaWorld 在長時程生成方面達到了最佳表現。作為一個全棧、開源且長期的專案,AlayaWorld 旨在為未來互動式影片世界模型的研究提供一個可擴展的基礎。
English
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.