ChatPaper.aiChatPaper

線形アテンションアーキテクチャ:メカニズム、トレードオフ、およびクロスレイヤールーティング

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

July 8, 2026
著者: Tommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag
cs.AI

要旨

自己注意機構により各トークンは全コンテキストから情報を取得できるが、系列長に対する二次コストによって長いコンテキストでの学習と推論が制限される。本稿では、ソフトマックス注意と、4つの最近のリカレント線形注意アーキテクチャ(DeltaNet、Gated DeltaNet、Kimi Delta Attention、Gated DeltaNet-2)の比較研究を提示する。これらの機構を共通のリカレントメモリ表記で表現し、表現力、メモリ減衰、消去および書き込み制御、トレーニングスループット、実装の複雑さにおいてどのように異なるかを明示する。実験は、150億トークンで学習された3億5000万パラメータモデルを中心とし、オプティマイザと学習率の比較、ハイブリッドスタックと純粋スタックの比較、系列長の実行時間測定、13億および30億パラメータでのより大規模なDeltaNet実行、および少数の下流評価を含む。報告される速度結果は、トレーニングスループットとイテレーション時間を測定したものであり、推論速度の実証的なベンチマークは提供しない。報告された3億5000万パラメータ、150億トークンのスイープにおいて、Muonを用いたKimi Delta Attentionが最も低い最終検証損失に達し、AdamWで学習された純粋なGated DeltaNetスタックが最も高い正規化トレーニングスループットを持ち、ハイブリッドスタックは一般にスループットのコストで損失を改善し、Muonは評価した一致したアーキテクチャ設定においてAdamWと比較して一貫して最終検証損失を低減する。我々は、DeltaNetスタイルのメモリに対する軽量なクロス層ルーティング機構を導入し評価する。最も自然なDeltaNetに着想を得た定式化、すなわち下層のデルタルール書き込み誤差を次の層の値ターゲットに転送する方法は、一致したベースラインを上回らない。代わりに整列された隠れストリームへのルーティングと書き込み値の転送は、報告する一致した実行において控えめな改善をもたらす。クロス層バリュールーティング(CLVR)は、DeltaNetとGated DeltaNetの両方で最終検証損失を低減する。
English
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.