고전적 캐시 정책이 실패할 때: 의미론적 검색 버퍼를 위한 학습 증강 교체
When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers
July 1, 2026
저자: Yushi Sun, Bowen Cao, Wai Lam
cs.AI
초록
LLM 에이전트는 과거 경험을 저장하고 재사용하기 위해 점차 검색 버퍼에 의존하고 있지만, 이러한 버퍼를 관리하는 캐시 관리 정책은 대부분 임시방편적인 수준에 머물러 있다. 우리는 이 문제를 전환 비용이 포함된 온라인 의미론적 캐시 교체 문제로 공식화한다. 여기서 항목은 임베딩 유사도로 매칭되며, 적중 품질은 이진(binary)이 아닌 연속적인 값을 가진다. MemoryBench-Full의 두 데이터셋(LoCoMo, DialSim)에 대해 8가지 교체 정책을 실험한 결과, 놀라운 사실을 발견했다: 고전적 휴리스틱(LRU, LFU)은 시간적 지역성(temporal locality)과 빈도 집중도(frequency concentration)가 부재하기 때문에 의미론적 작업 부하에서 단순한 FIFO 기준선보다 일관되게 낮은 성능을 보인다. 우리는 학습 기반 프레임워크인 SOLAR를 제안한다. 이 프레임워크는 후회 누적(regret accumulation)으로부터 수정 시점을 도출하고(약 17%의 수정률 달성), 암시적 검색 피드백에 대한 베이지안 온라인 학습(Bayesian online learning)을 통해 내용 선택을 수행한다. 우리는 SOLAR가 캐시 크기와 시간 지평(horizon)에 독립적인 상수 경쟁비(competitive ratio) ≤ 3을 달성함을 증명한다(FIFO의 Ω(K)와 대조적). 또한 퇴출 후회(eviction regret)가 O(KT log T)임을 보여주며, 이는 Ω(KT) 하한을 로그 인자까지 일치시킨다. 실험 결과, 빡빡한 캐시 크기에서 FIFO 대비 5~75%의 상대적 개선을 보였으며, 작업 집합(working set) 경계에서 명확히 특성화된 상전이(phase transition)가 관찰되었다. 5000개 항목 풀(pool)을 사용한 합성 실험에서는 풀 크기와 검색 품질 간의 역U자형 관계(inverted-U relationship)가 추가로 드러났으며, 이는 용량 제약이 저장 한계가 아닌 검색 잡음 현상(retrieval noise phenomenon)임을 입증한다.
English
LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain largely ad-hoc. We formalize this as an online semantic cache replacement problem with switching costs, where items are matched by embedding similarity and hit quality is continuous rather than binary. Through experiments on two datasets from MemoryBench-Full (LoCoMo, DialSim) with 8 replacement policies, we reveal a surprising finding: classic heuristics (LRU, LFU) consistently underperform the naive FIFO baseline on semantic workloads, due to the absence of temporal locality and frequency concentration. We propose SOLAR, a learning-augmented framework that derives modification timing from regret accumulation (achieving sim17\% modification rate) and content selection from Bayesian online learning over implicit retrieval feedback. We prove SOLAR achieves a constant competitive ratio leq 3, independent of cache size and horizon (vs.\ Ω(K) for FIFO), and eviction regret O(KTlog T), matching the Ω(KT) lower bound up to logarithmic factors. Experiments demonstrate 5--75\% relative improvement over FIFO at tight cache sizes, with a clearly characterized phase transition at the working set boundary. Synthetic experiments with 5000-item pools further reveal an inverted-U relationship between pool size and retrieval quality, justifying capacity constraints as a retrieval noise phenomenon rather than a storage limitation.