MemHarness: 메모리는 재생되는 것이 아니라 재구성된다
MemHarness: Memory Is Reconstructed, Not Replayed
July 30, 2026
저자: Rong Wu, Daocheng Fu, Licheng Wen, Xuemeng Yang, Shu Zou, Jianbiao Mei, Yuxin Wang, Hairong Zhang, Yu Yang, Tao Hu, Cong Zhang, Botian Shi, Pinlong Cai
cs.AI
초록
과거 경험을 검색하는 것은 대규모 언어 모델 에이전트의 성능을 향상시키기 위한 일반적인 전략이 되었다. 그러나 기존의 대부분 메모리 증강 에이전트는 검색된 경험을 문자 그대로 재생할 정적 기록으로 취급하여, 에이전트의 현재 상황과 부합하는지와 무관하게 이를 맥락에 주입한다. 이러한 "재생" 패러다임은 저장된 경험의 추상적이고 일반적인 특성과 의사결정 시점에 직면하는 구체적이고 끊임없이 변화하는 상태 사이의 간극을 무시하며, 흔히 부정적 전이를 초래한다. 반면, 인간은 과거 경험을 문자 그대로 회상하는 경우가 드물다. 대신, 인간은 검색된 기억을 현재 맥락에 맞게 재구성하고 적응시킨다. 이에 영감을 받아, 우리는 LLM 에이전트가 현재 맥락에 기반하여 과거 경험을 능동적으로 활용하고 재구성하도록 하는 프레임워크인 MemHarness를 제안한다. 각 의사결정 단계에서 통합 정책 모델은 현재 상태에 기반하여 검색된 경험을 비판하고 재구성하며, 행동 이전에 맥락에 근거한 지침을 생성한다. 이러한 재구성 능력은 GRPO를 통한 종단 간 훈련을 통해 자연스럽게 발현된다. ALFWorld와 WebShop에서의 실험은 MemHarness가 순수 RL 및 정적 메모리 증강 기준선보다 현저히 우수한 성능을 보여주며, 분포 외(OOD) 시나리오에서 강력한 견고성을 입증한다. 또한, 우리의 분석은 이러한 재구성 목적이 부정적 전이를 방지할 뿐만 아니라 훈련 중 잠재적 지침으로 작용하여 에이전트의 내재적 추론 능력을 근본적으로 향상시킴을 보여준다.
English
Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent's intrinsic reasoning capabilities.