铭记于心:智能体记忆中内隐联想盲点的基准测试
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
July 27, 2026
作者: Ruizhe Li, Mingxuan Du, Benfeng Xu, Zhendong Mao
cs.AI
摘要
长期记忆系统会将用户所述内容存储于外部存储器中,并在相关查询出现时予以检索。这一接口建立在一条看似理所当然却鲜少明言的假设之上:所需记忆必然与需要它的查询存在相似性。世界知识打破了这一假设。树坚果过敏应通过杏仁粉这一成分来改变对马卡龙请求的回答,然而两段文本之间并无检索器可识别的线索。我们将这种失效模式称为"隐含关联盲点",并提出了InMind——一个包含125项任务、经专家验证的基准测试集,覆盖十个生活领域,其中113项任务基于可引用的公开来源。其配对对照组分离了现有评估中混淆的三种解释:事实从未存储、模型缺乏关联知识、或事实已存储却未能浮现。结论清晰明确。当决定性的记忆被置于上下文中时,骨干模型能正确回答84.0%的间接查询;而当同一记忆必须通过检索获取时,六种向量、图与智能体记忆系统最高仅能达到14.4%,尽管它们在有需求时对相同事实的召回率可达100%。将嵌入维度提升八倍后,每个系统在答案无关的目标召回率上均有提高,但差距基本保持不变。一种在查询到来前保持记忆可见的最小诊断探针能弥合大部分差距,将失效根源定位在查询条件化接口本身,并指向路由选择——即决定哪些事实必须保持可见——作为InMind旨在评估的开放性问题。
English
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.