ChatPaper.aiChatPaper

AtlasVLA: 비전-언어-행동 모델을 위한 지속적 세계-자아 상태 모델링

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

August 7, 2026
저자: Guiyu Zhao, Longteng Guo, Yanghong Mei, Zilin Zhu, Yu Zhang, Bin Cao, Mingming Yu, Xingjian He, Jie Jiang, Jing Liu
cs.AI

초록

비전-언어-행동(VLA) 모델이 구현형 AI를 발전시켰지만, 이들의 근본적으로 반응적인 패러다임은 부분 관측 가능하고 장기적인 과제에서 성능을 심각하게 제한한다. 단일 손목 장착 카메라로 제한될 때, 이러한 모델은 객체가 시야를 벗어남에 따라 지각 망각과 다단계 실행 중 시간적 작업 진행 망각을 필연적으로 겪는다. 이러한 병목을 극복하기 위해, 우리는 직접적인 반응적 조작에서 지속적인 세계-자아 상태를 통한 능동적 추론으로 전환하는 새로운 프레임워크인 AtlasVLA를 제안한다. AtlasVLA는 이중 메모리 아키텍처를 특징으로 한다: 일시적인 2D 관측을 전역적으로 갱신되는 복셀 해시 공간 상태로 승격시켜 시각적 사각지대를 해결하는 4D 지속 세계 상태 메모리(4D Persistent World State Memory)와, 과거 자아 상태와 작업 진행 상황을 추적하는 자아 작업 상태 메모리(Ego-Working State Memory)가 그것이다. 결합된 세계-자아 상태에 확산 트랜스포머(DiT)를 조건화함으로써 AtlasVLA는 강건한 공간 추론을 가능하게 한다. LIBERO, RLBench 및 실제 환경 벤치마크에 걸친 광범위한 평가는 AtlasVLA가 손목 카메라만을 사용하여 최첨단 성능을 달성함을 입증한다. 놀랍게도, 이는 다중 뷰 베이스라인을 압도적으로 능가하며, LIBERO-Long에서 9.4%, 실제 환경 장기 과제에서 17.5%의 절대적 성공률 향상을 달성한다.
English
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.