銘記在心:智能體記憶中內隱聯想盲點的基準測試
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
July 27, 2026
作者: Ruizhe Li, Mingxuan Du, Benfeng Xu, Zhendong Mao
cs.AI
摘要
長期記憶系統將使用者所言儲存於外部儲存器中,並在相關查詢到來時將其檢索出來。此類介面基於一個極其自然以致鮮少明言的假設:所需的記憶將與需要該記憶的查詢相似。世界知識打破了這一假設。樹堅果過敏本應透過馬卡龍中的杏仁粉成分改變對馬卡龍請求的回答,然而這兩段文字之間並無檢索器可見的線索。我們將此種失敗模式稱為「隱含關聯盲點」,並引入InMind——一個涵蓋125項任務、經專家驗證的基準測試,橫跨十個生活領域,其中113項任務奠基於可引用的公開來源。其配對對照組將現有評估中混淆的三種解釋區分開來:事實從未被儲存、模型缺乏橋接知識、或事實已被儲存但從未浮現。結論清晰明確。當關鍵記憶置於上下文中時,骨幹模型能回答84.0%的間接查詢;而當同一記憶必須被檢索出來時,六個向量、圖及智能體記憶系統的表現最多僅達14.4%,儘管它們按需回憶相同事實的準確率可達100%。嵌入維度提升八倍雖能提高每個系統的無答案目標召回率,但差距基本保持不變。一種在查詢到達前保持記憶可見的最小診斷探針,恢復了大部分差距,將失敗定位於查詢條件介面本身,並指向路由——決定哪些事實必須保持可見——作為InMind旨在評分的開放問題。
English
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.