AtlasVLA:視覚・言語・行動モデルのための持続的世界・自己状態モデリング
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
August 7, 2026
著者: Guiyu Zhao, Longteng Guo, Yanghong Mei, Zilin Zhu, Yu Zhang, Bin Cao, Mingming Yu, Xingjian He, Jie Jiang, Jing Liu
cs.AI
要旨
Vision-Language-Action(VLA)モデルは具現化AIを大きく前進させてきたが、その根本的に反応的なパラダイムは、部分観測タスクや長期的タスクにおける性能を深刻に制限している。単一の手首搭載カメラに限定された場合、物体が視野から外れるにつれて生じる知覚の忘却と、多段階実行中に生じる時間的タスク進捗の忘却を不可避的に引き起こす。これらのボトルネックを克服するため、我々はAtlasVLAを提案する。これは、直接的な反応的操作から、持続的なワールド・エゴ状態に基づく能動的推論へと転換する新規フレームワークである。AtlasVLAは二重メモリアーキテクチャを特徴とする。すなわち、一時的な2次元観測を大域的に更新されるボクセルハッシュ空間状態へと変換し、視覚的な死角を解消する4D持続的世界状態メモリと、過去のエゴ状態とタスク進捗を追跡するエゴ作業状態メモリである。この統合されたワールド・エゴ状態を用いて拡散トランスフォーマー(DiT)を条件付けることにより、AtlasVLAは頑健な空間推論を実現する。LIBERO、RLBench、および実世界ベンチマークにわたる広範な評価により、AtlasVLAが手首カメラのみを用いて最先端の性能を達成することが実証された。特筆すべきことに、AtlasVLAは多視点ベースラインを決定的に上回り、LIBERO-Longでは9.4%、実世界の長期的タスクでは17.5%の絶対成功率の向上を達成した。
English
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.