ChatPaper.aiChatPaper

선형 어텐션 아키텍처: 메커니즘, 트레이드오프, 그리고 계층 간 라우팅

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

July 8, 2026
저자: Tommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag
cs.AI

초록

자기 주의(self-attention)는 각 토큰이 전체 맥락에서 정보를 가져올 수 있게 하지만, 시퀀스 길이에 따른 제곱 비용(quadratic cost)이 장기 맥락에서의 학습과 추론을 제한한다. 본 논문은 소프트맥스 주의(softmax attention)와 네 가지 최신 순환 선형 주의 아키텍처인 DeltaNet, Gated DeltaNet, Kimi Delta Attention, Gated DeltaNet-2를 비교 연구한다. 이러한 메커니즘을 공통된 순환 메모리 표기법(recurrent-memory notation)으로 표현하여, 표현력(expressivity), 메모리 감쇠(memory decay), 삭제 및 쓰기 제어(erase and write control), 학습 처리량(training throughput), 구현 복잡성(implementation complexity)에서 어떻게 차이가 나는지 명시적으로 보여준다. 실험은 150억 토큰으로 학습된 3억 5천만 파라미터(350M-parameter) 모델을 중심으로 진행되었으며, 최적화기 및 학습률 비교, 하이브리드 대 순수 스택 비교, 시퀀스 길이별 실행 시간 측정, 13억 및 30억 파라미터 규모의 더 큰 DeltaNet 실험, 그리고 소규모 하류 평가(downstream evaluations)를 포함한다. 보고된 속도 결과는 학습 처리량과 반복 시간을 측정한 것이며, 추론 속도 벤치마크(inference-speed benchmark)는 별도로 제공하지 않는다. 보고된 3억 5천만 파라미터 및 150억 토큰 실험 범위 내에서 Muon을 사용한 Kimi Delta Attention이 가장 낮은 최종 검증 손실에 도달했고, AdamW로 학습된 순수 Gated DeltaNet 스택이 가장 높은 정규화된 학습 처리량을 보였으며, 하이브리드 스택은 일반적으로 처리량 비용을 감수하면서 손실을 개선했고, Muon은 일치된 아키텍처 설정에서 평가한 모든 경우에 AdamW에 비해 최종 검증 손실을 일관되게 낮췄다. 또한 DeltaNet 스타일 메모리를 위한 경량 교차 계층 라우팅 메커니즘(cross-layer routing mechanisms)을 도입하고 평가한다. 가장 자연스러운 DeltaNet 기반 공식(하위 계층의 델타 규칙 쓰기 오류(delta-rule write error)를 다음 계층의 값 타겟(value target)에 전달)은 일치된 기준선 대비 개선되지 않았다. 정렬된 은닉 스트림(aligned hidden stream)으로 라우팅하고 쓰기 값(write value)을 대신 전달하는 방식은 보고된 일치된 실험에서 약간의 개선을 보였다: 교차 계층 값 라우팅(Cross-Layer Value Routing, CLVR)은 DeltaNet과 Gated DeltaNet 모두에서 최종 검증 손실을 낮췄다.
English
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.