ABot-AgentOS:一种具备终身多模态记忆的通用机器人智能体操作系统

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

July 11, 2026
作者: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu
cs.AI

摘要

近期视觉语言模型(VLM)与视觉语言动作(VLA)系统提升了机器人感知与动作预测能力,但长时域具身智能体仍需通用运行时层来支持推理、记忆、工具使用、验证及跨本体执行。我们提出ABot-AgentOS——一种通用机器人代理操作系统,位于底层控制器之上,提供面向场景规划的审慎代理层、上下文隔离的技能执行、多阶段验证、多模态记忆及边云协同。为评估此类系统,我们引入EmbodiedWorldBench——一个包含16个室内外及混合场景、四个难度等级、涵盖导航、物体搜索、NPC对话、动态事件及基于轨迹评分的200余项任务的可执行基准测试。ABot-AgentOS进一步提出通用多模态图记忆(Universal Multi-modal Graph Memory),这是一种持久化、源锚定的基板,可将对话、视觉观测、空间上下文、时序关系及任务轨迹转化为类型化节点与边。其故障驱动的自我进化循环可将诊断出的记忆故障转化为门控运行时进化资产,仅后续评估分割可访问,既避免当前分割的真相泄露,又实现持续改进。在EmbodiedWorldBench初始子集上,ABot-AgentOS在任务成功率与目标完成度上均优于单控制器基线。跨记忆基准测试中,ABot-AgentOS静态版本在LoCoMo上达87.5,在OpenEQA EM-EQA上达59.9,在Mem-Gallery上达88.6,在NExT-QA上以76.5的整体准确率领先;自我进化进一步将LoCoMo提升至88.7,OpenEQA提升至60.4,Mem-Gallery提升至89.0。结果表明,通用Agent OS层可提升长时域具身执行能力,并为持续交互提供持久、可审计的记忆。
English
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.
PDF681July 15, 2026