ABot-AgentOS: 生涯持続型マルチモーダルメモリを備えた汎用ロボットエージェントOS
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
July 11, 2026
著者: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu
cs.AI
要旨
近年のVLMおよびVLAシステムにより、ロボットの知覚と行動予測は向上しましたが、長期的なタスクを遂行する身体化エージェントには、推論、記憶、ツール使用、検証、異なる身体での実行を可能にする汎用ランタイム層が依然として必要です。本稿では、低レベル制御装置の上位に位置し、シーン条件付き計画、コンテキスト分離型スキル実行、多段階検証、マルチモーダル記憶、エッジ・クラウド連携を提供する、汎用ロボットエージェントオペレーティングシステム「ABot-AgentOS」を提案します。このようなシステムを評価するため、16の屋内・屋外・ハイブリッドシーン、4段階の難易度、ナビゲーション・物体探索・NPC対話・動的イベント・トレースに基づく評価を含む200以上のタスクで構成される実行可能なベンチマーク「EmbodiedWorldBench」を導入します。ABot-AgentOSはさらに、対話、視覚観測、空間コンテキスト、時間関係、タスクトレースを型付きノードとエッジに変換する、永続的なソースに基づいた基盤「Universal Multi-modal Graph Memory」を導入します。失敗駆動型自己進化ループは、診断された記憶の失敗を、ランタイムでゲート制御された進化資産に変換し、これらは後続の評価分割に対してのみ昇格されるため、現在の分割における正解データの漏洩を防ぎつつ、継続的な改善を可能にします。初期のEmbodiedWorldBenchサブセットにおいて、ABot-AgentOSは単一コントローラベースラインと比較して、タスク成功率と目標達成度の両方で改善を示しました。記憶ベンチマークでは、ABot-AgentOS StaticはLoCoMoで87.5、OpenEQA EM-EQAで59.9、Mem-Galleryで88.6、NExT-QAでAcc@All 76.5を達成し、自己進化によりLoCoMoは88.7、OpenEQAは60.4、Mem-Galleryは89.0にさらに向上しました。これらの結果は、汎用Agent OS層が長期的な身体化実行を改善し、継続的な対話のための永続的で監査可能な記憶を提供できることを示唆しています。
English
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.