칼만 델타 네트워크: 불확실성 인식 연상 메모리
Kalman Delta Networks: Uncertainty-aware Associative Memory
September 7, 2026
저자: Ngoc Bui, Tinglin Huang, Rex Ying
cs.AI
초록
선형 어텐션은 효율적인 장문맥 추론과 상수 메모리 디코딩을 위해 최첨단 언어 모델에서 점점 더 많이 사용되고 있다. 그러나 고정 크기 순환 메모리는 각 토큰마다 온라인 결정을 요구한다: 미래의 질의가 어떤 정보를 필요로 할지 알기 전에 무엇을 기록할지, 기존 연상을 얼마나 강하게 덮어쓸지 말이다. 델타 규칙 모델은 이 강도를 현재 토큰 임베딩으로부터 학습하지만 메모리 추정에 대한 신뢰도를 추적하지 않아, 각 쓰기가 축적된 증거에 적응하는 것을 막는다. 이러한 불확실성을 명시적으로 표현하기 위해 우리는 순환 연상 메모리를 칼만 필터가 최적의 재귀 추정기인 선형-가우시안 상태공간 모델로 재정식화하고, 새로운 모델 계열인 칼만 델타 네트워크(Kalman Delta Networks, KDN)를 도입한다. KDN에서 전이는 메모리 상태와 그 불확실성을 모두 전파하여, 칼만 이득이 각 잔차 쓰기를 축적된 증거와 관측 신뢰도에 따라 가중하도록 한다. 이 정식화 아래에서 델타 스타일 업데이트는 예측 공분산에 대한 토큰별 등방성 대체물을 대입하고 공분산 추적을 생략하는 특수 사례로 나타난다. 그러나 정확한 추적은 GPU 병렬 선형 어텐션 스캔에 적합하지 않은 조밀한 상태 의존적 리카티 재귀를 수반한다. 이 문제를 해결하기 위해 우리는 스캔 호환 KDN 근사 두 가지를 도입한다. 대각 KDN은 온라인 평균장 변분 추론을 통해 각 1단계 사후분포를 대각 가우시안 계열로 사영하는 반면, 등방성 KDN은 헤드당 단일 불확실성 스칼라를 갖는 등방성 근사를 사용한다. 이들의 불확실성 재귀는 뫼비우스 사상이며, 로그 병렬 깊이의 결합 스캔을 가능하게 한다. 750M 및 1.3B 파라미터에서의 통제된 사전학습 전반에 걸쳐, KDN 변형들은 최신 선형 어텐션 모델보다 퍼플렉서티와 평균 다운스트림 정확도를 일관되게 개선한다.
English
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.