대규모 메모리 디코더: 사전 훈련된 파라메트릭 장기 메모리
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
July 30, 2026
저자: Rubin Wei, Jiaqi Cao, Jiarui Wang, Junming Zhang, Qipeng Guo, Bowen Zhou, Zhouhan Lin
cs.AI
초록
디코더 전용 언어 모델은 장기 메모리와 추론을 단일 파라미터 집합에 얽어매므로 메모리 용량을 독립적으로 확장하기 어렵게 만든다. Memory Decoder는 파라메트릭 장기 메모리 모듈을 도입하지만, 비교적 작은 규모에서만 이를 연구했다. 본 연구에서는 메모리 모델을 최대 6.9B 파라미터로 확장하고 300B 토큰으로 사전학습한 Memory Decoder at Scale을 제시한다. 이러한 데이터 규모에서 인덱싱과 검색의 결합 비용은 표준 Faiss 파이프라인을 실행 불가능하게 만든다. 우리는 Faiss 인덱싱 및 검색을 위한 분산 파이프라인과 kNN 분포의 희소 배치 단위 로딩을 통해 이 병목 문제를 해결한다. 모델 규모 전반에 걸쳐, 메모리에 더 많은 파라미터를 할당하는 것이 기본 모델만 확장하는 것보다 더 나은 파라미터-성능 트레이드오프를 제공함을 확인했다. 17개 벤치마크에서 6.9B 일반 메모리를 Pythia-410M과 짝지으면 평균 점수가 29.86에서 37.34로 올라가며, 총 파라미터가 39% 더 적으면서도 Pythia-12B(37.24)를 능가한다. 0.6B에서 14B 범위의 Qwen3 Base 모델의 경우, 1.7B 도메인 메모리는 모든 규모에서 세 도메인에 걸친 평균 점수를 9점 이상 향상시킨다. 전체적으로, 우리의 결과는 사전학습된 메모리를 독립적으로 확장하는 것이 언어 모델 성능을 향상시키는 더 파라미터 효율적인 경로를 제공함을 보여준다.
English
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.