ABot-World-0: 단일 데스크톱 GPU에서의 무한 인터랙티브 월드 롤아웃
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
July 21, 2026
저자: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
cs.AI
초록
본 논문에서는 실시간 장기 폐쇄 루프 상호작용을 위한 행동 조건부 비디오 월드 모델인 ABot-World-0을 제시합니다. 이 모델은 AAA 게임, 시뮬레이션 엔진, 인터넷 비디오를 포괄하는 다중 소스 데이터 인프라를 기반으로 제어 가능한 월드 다이내믹스를 학습합니다. WorldExplorer는 훈련 피드백에 따라 에이전트 기반 수집을 수행하며, 통합 파이프라인은 14개의 결정론적 품질 검사, VLM 기반 평가, 동기화된 행동 및 텍스트 주석을 적용합니다. 우리는 점진적으로 양방향 행동 조건부 교사 모델을 교사 강제(teacher forcing) 및 ODE 증류(ODE distillation)를 통해 인과적 학생 모델로 증류하며, 장기 학생 자기 롤아웃을 확장된 지평 교사와 정렬하기 위해 LongForcing을 도입하여 누적 분포 이동 및 자기회귀 드리프트를 완화합니다. 원시 키보드 동작은 장면 로밍 및 3인칭 캐릭터 상호작용을 위한 통합 제어 인터페이스를 제공하며, 참조 캐릭터 메모리는 3인칭 롤아웃 중 정체성 일관성을 위한 지속적인 외형 단서를 제공합니다. 배포를 위해 우리는 경량 VAE 디코더, 효율적인 어텐션, 메모리 인식 스케줄링, 저비트 DiT 추론을 갖춘 스트리밍 추론 스택을 공동 설계합니다. 최적화된 저비트 구성에서 ABot-World-0는 단일 NVIDIA RTX 5090 데스크탑 GPU에서 최대 16 FPS로 720P 비디오를 스트리밍하며, 동작-첫 프레임 지연 시간은 1.2초, 최대 VRAM 사용량은 약 19GiB입니다. WorldRoamBench 및 확장된 대화형 롤아웃에 대한 실험은 경쟁력 있는 제어 가능성과 일관된 장기 월드 진화를 입증합니다.
English
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.