ChatPaper.aiChatPaper

AtlasVLA:面向视觉-语言-动作模型的持久世界-本体状态建模

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

August 7, 2026
作者: Guiyu Zhao, Longteng Guo, Yanghong Mei, Zilin Zhu, Yu Zhang, Bin Cao, Mingming Yu, Xingjian He, Jie Jiang, Jing Liu
cs.AI

摘要

尽管视觉-语言-动作(VLA)模型推动了具身AI的发展,但其本质上的反应式范式严重制约了在部分可观察和长时程任务中的表现。当仅依赖单个腕部摄像头时,模型不可避免地会因物体移出视野而产生感知遗忘,并在多步执行过程中出现时序任务进度遗忘。为克服这些瓶颈,我们提出AtlasVLA——一种全新框架,通过持久的“世界-自我”状态实现从直接反应式操作向主动推理的转变。AtlasVLA采用双记忆架构:4D持久世界状态记忆将瞬时的2D观测提升为全局更新的体素哈希空间状态,以解决视觉盲区问题;自我工作状态记忆则跟踪历史自我状态与任务进度。通过以联合的“世界-自我”状态为条件引导扩散变换器(DiT),AtlasVLA实现了鲁棒的空间推理。在LIBERO、RLBench及真实世界基准上的广泛评估表明,AtlasVLA仅使用腕部摄像头即可达到最先进的性能。值得注意的是,它显著优于多视角基线方法,在LIBERO-Long上实现了9.4%的绝对成功率提升,在真实世界长时程任务中实现了17.5%的提升。
English
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.