ChatPaper.aiChatPaper

기억하세요: 에이전트 메모리에서의 암묵적 연관성 맹점 벤치마킹

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

July 27, 2026
저자: Ruizhe Li, Mingxuan Du, Benfeng Xu, Zhendong Mao
cs.AI

초록

장기 기억 시스템은 사용자의 발화를 외부 저장소에 저장하고, 관련 질의가 도착하면 이를 검색한다. 이 인터페이스는 거의 명시적으로 언급되지 않을 정도로 자연스러운 가정에 기반한다: 필요한 기억은 그것을 필요로 하는 질의와 유사할 것이라는 가정이다. 세계 지식은 이 가정을 깨뜨린다. 견과류 알레르기가 있다면 마카롱 요청에 대한 답변이 아몬드 가루 성분을 통해 바뀌어야 하지만, 두 텍스트 사이에는 검색기가 인식할 수 있는 단서가 전혀 없다. 우리는 이러한 실패 모드를 '암묵적 연관성 사각지대'라고 명명하고, 열 개의 생활 영역에 걸쳐 125개 과제로 구성된 전문가 검증 벤치마크인 InMind를 소개한다. 이 중 113개 과제는 인용 가능한 공개 출처에 기반한다. 이 벤치마크의 짝지어진 대조군은 기존 평가에서 혼동되던 세 가지 설명을 분리한다: 사실이 저장되지 않은 경우, 모델이 연결 지식을 결여한 경우, 또는 사실이 저장되었으나 표면화되지 않은 경우. 결론은 명확하다. 결정적 기억이 문맥에 주어졌을 때 백본은 간접 질의의 84.0%를 답한다. 동일한 기억을 검색해야 할 때, 여섯 가지 벡터, 그래프 및 에이전트 메모리 시스템은 최대 14.4%에 도달하며, 이는 이들 시스템이 요청 시 동일한 사실을 최대 100% 회상함에도 불구하고 그렇다. 차원이 8배 더 큰 임베딩은 모든 시스템에서 답변-블라인드 대상 회상을 향상시키지만, 격차는 본질적으로 그대로 유지된다. 질의 도착 전에 기억을 가시적으로 유지하는 최소 진단 프로브는 격차의 대부분을 회복하며, 실패의 위치를 질의-조건화 인터페이스 자체로 좁히고, 어떤 사실을 가시적으로 유지할지 결정하는 라우팅이 InMind가 평가하도록 설계된 미해결 문제임을 지적한다.
English
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.