ABot-AgentOS: 평생 멀티모달 메모리를 갖춘 범용 로봇 에이전트 OS
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
July 11, 2026
저자: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu
cs.AI
초록
최근 VLM 및 VLA 시스템은 로봇의 인지 및 행동 예측 능력을 향상시켰지만, 장기적 행동을 수행하는 내장형 에이전트는 여전히 추론, 메모리, 도구 사용, 검증 및 교차 개체 실행을 위한 일반 런타임 계층을 필요로 합니다. 본 논문에서는 저수준 제어기 위에 위치하며 장면 조건부 계획, 맥락 분리형 기술 실행, 다단계 검증, 다중 모달 메모리, 에지-클라우드 협업을 위한 숙고형 에이전트 계층을 제공하는 일반 로봇 에이전트 운영체제인 ABot-AgentOS를 제시합니다. 이러한 시스템을 평가하기 위해, 16개의 실내, 실외 및 혼합 장면, 4개의 난이도 수준, 그리고 탐색, 객체 탐색, NPC 대화, 동적 이벤트 및 트레이스 기반 점수화를 포함한 200개 이상의 태스크로 구성된 실행 가능한 벤치마크인 EmbodiedWorldBench를 소개합니다. ABot-AgentOS는 또한 대화, 시각적 관찰, 공간적 맥락, 시간적 관계 및 태스크 트레이스를 유형화된 노드와 엣지로 변환하는 지속적인 소스 기반 기질(substrate)인 범용 다중 모달 그래프 메모리(Universal Multi-modal Graph Memory)를 도입합니다. 오류 기반 자가 진화 루프는 진단된 메모리 오류를 게이트된 런타임 진화 자산(gated runtime evo-assets)으로 변환하여, 이후 평가 분할에서만 승격시킴으로써 현재 분할의 정답 유출을 방지하면서 지속적 개선을 가능하게 합니다. EmbodiedWorldBench 초기 하위 집합에서 ABot-AgentOS는 단일 컨트롤러 기준선 대비 태스크 성공률과 목표 완료율 모두에서 향상된 성능을 보였습니다. 메모리 벤치마크 전반에서 ABot-AgentOS Static은 LoCoMo에서 87.5, OpenEQA EM-EQA에서 59.9, Mem-Gallery에서 88.6, NExT-QA에서 76.5의 Acc@All을 달성했으며, 자가 진화를 통해 LoCoMo는 88.7, OpenEQA는 60.4, Mem-Gallery는 89.0으로 추가 향상되었습니다. 이러한 결과는 일반 에이전트 OS 계층이 지속적 상호작용을 위한 영구적이고 감사 가능한 메모리를 제공하면서 장기적 내장형 실행을 개선할 수 있음을 시사합니다.
English
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.