ABot-AgentOS:一個具備終身多模態記憶的通用機器人智能體操作系統

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

July 11, 2026
作者: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu
cs.AI

摘要

近期的VLM與VLA系統已顯著提升機器人的感知與動作預測能力,然而長時域具身智能體仍需一套通用的運行時層,以支援推理、記憶、工具使用、驗證及跨本體執行。我們提出ABot-AgentOS,這是一套通用的機器人智慧體作業系統,位於低階控制器之上,提供一個慎思式智慧體層,實現場景條件規劃、上下文隔離的技能執行、多階段驗證、多模態記憶以及邊緣-雲端協作。為評估此類系統,我們引入EmbodiedWorldBench,這是一個可執行的基準測試,涵蓋16個室內、室外及混合場景,四個難度等級,以及超過200個任務,包含導航、物體搜尋、NPC對話、動態事件與軌跡導向評分。ABot-AgentOS更進一步引入通用多模態圖形記憶,這是一個持久性的源導向基底,可將對話、視覺觀測、空間脈絡、時間關係與任務軌跡轉化為具類型的節點與邊。一套失敗驅動的自我演化循環會將診斷出的記憶缺陷轉化為門控運行時進化資產,這些資產僅在後續的評估拆分中提升,從而防止當前拆分中的真值洩漏,同時實現持續改進。在EmbodiedWorldBench的初始子集上,ABot-AgentOS在任務成功率與目標完成度上均優於單一控制器基線。在記憶基準測試中,ABot-AgentOS靜態版於LoCoMo達到87.5,於OpenEQA EM-EQA達到59.9,於Mem-Gallery達到88.6,於NExT-QA達到76.5的Acc@All;自我演化進一步將LoCoMo提升至88.7,OpenEQA提升至60.4,Mem-Gallery提升至89.0。這些結果表明,通用智慧體作業系統層能夠改善長時域具身執行,同時為持續互動提供持久且可稽核的記憶。
English
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.
PDF681July 15, 2026