ReflectWorld-MM: 개방형 비디오 스트림을 위한 개체 중심의 다중 모달 메모리 시스템
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
July 14, 2026
저자: Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
cs.AI
초록
세상을 지속적으로 관찰하고, 본 것을 기억하며, 축적된 경험을 바탕으로 추론할 수 있는 어시스턴트를 구축하는 것은 오랜 목표였으며, 최근에는 비디오 스트림에 대한 장기 기억을 갖춘 멀티모달 에이전트가 점점 더 많은 관심을 받고 있습니다. 불행히도, 기존 시스템은 메모리를 모델 컨텍스트 내부나 평면적인 특징 저장소에 보관하며, 스트림이 실제로 다루는 지속적인 개체가 아닌 프레임 중심으로 구성합니다. 이는 시스템을 제한된 동영상에 국한시키고, 시간이 지남에 따라 누가 무엇이 다시 나타나는지 추적하는 능력을 약화시킵니다. 본 논문에서는 개방형 비디오 스트림을 위한 개체 중심의 멀티모달 메모리 시스템인 ReflectWorld-MM을 제안합니다. 이는 세 부분으로 구성됩니다. 첫 번째는 인식 프론트엔드로, 제한된 단기 기억 하에서 시청각 스트림을 개체로 분해된 관찰로 변환합니다. 두 번째는 인간 기억 이론에 기반한 계층적 장기 기억으로, 다중 스케일의 일화적 기억, 진화하는 개체 중심의 의미 기억, 그리고 절차적 기억을 결합합니다. 세 번째는 실제 환경에서 작동하도록 구축된 완전한 구현체로, 임의의 스트림을 입력받아 기성 어시스턴트에 연결됩니다. 여섯 개의 장편 비디오 및 평생 기억 벤치마크에서 ReflectWorld-MM은 모든 여섯 개에서 최고 정확도를 달성하여, 강력한 메모리 에이전트와 최첨단 모델을 능가했습니다.
English
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to bounded videos and weakens their ability to track who and what reappears over time. In this paper, we propose ReflectWorld-MM, an entity-oriented multimodal memory system for open-ended video streams. It consists of three parts. The first is a perception front-end that turns an audiovisual stream into entity-resolved observations under a bounded short-term memory. The second is a hierarchical long-term memory, grounded in human memory theory, that couples a multi-scale episodic memory, an evolving entity-centric semantic memory, and a procedural memory. The third is a complete realization, built for real-world operation, that ingests arbitrary streams and plugs into off-the-shelf assistants. Across six long-video and lifelong-memory benchmarks, ReflectWorld-MM achieves the best accuracy on all six, outperforming strong memory agents and a frontier model.