大型語言模型的記憶
Memory for Large Language Models
July 28, 2026
作者: Sining Zhoubian, Dan Zhang, Evgeny Kharlamov, Jie Tang
cs.AI
摘要
記憶已演變為大型語言模型(LLMs)中一個基礎性的架構維度,從計算過程中隱性的副產品,轉變為一系列明確且可操控的機制。儘管近期進展引入了多元策略——涵蓋短期注意力、遞迴狀態動態、參數高效調整與可擴展查找儲存——但此快速發展亦導致研究領域高度破碎化。本調查報告提出一套系統性、以架構為核心的LLMs記憶分類法。我們的分析框架沿三個正交軸向描述記憶:表徵(隱性 vs. 顯性)、更新動態(離線 vs. 在線)與持續性(短期 vs. 長期)。我們進一步形式化記憶寫入、路由、狀態轉換與整合的細粒度機制。此統一視角闡明了計算耦合記憶與獨立定址記憶之間的概念界線,有效銜接了不同架構典範。此外,我們批判性分析了混合記憶架構、系統層級效率權衡與多維度評估方法。透過將這些分散的進展整合為連貫框架,本調查報告描繪了以記憶為中心的LLM設計軌跡,並為未來可擴展且具適應性的語言模型創新提供了原則性基礎。
English
Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.