ChatPaper.aiChatPaper

べき乗則グラフアテンション:スケールド・ドット積アテンションの厳密な一般化、推論時における経験的崩壊

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

August 10, 2026
著者: Burc Gokden
cs.AI

要旨

Power Law Decoder Representations(PLDR-LLM)およびその注意機構であるPower Law Graph Attention(PLGA)は、スケールド・ドットプロダクト注意機構(SDPA)の固定された双線形形式を、正のテンソルA_{LM}から要素ごとの冪則によって構成される、学習され入力から生成される双線形作用素G_{LM}で置き換える。本アーキテクチャは完全に仕様化され、固定されたリファレンスリリースに対して検証されている。主張は、定理、条件付き定理、測定結果、または予想として分類される。 無条件に成り立つこと:PLGAは、G_{LM}=Iのとき正確にSDPAを含む。A_{LM}とA_Pはすべての成分について狭義正であり、A_{LM}はペロン=フロベニウス構造を持つ。DAG正則化項はNOTEARSのウォーク計数形式を持ち、正値性が厳密な非巡回性を妨げる。さらに、非共鳴条件(標準的なロータリー周波数によって満たされる)の下では、交換子(コミュタント)判定条件により、どの作用素が相対位置依存性を保つかが特定される。 推論崩壊定理:演繹的出力の厳密な入力不変性は、推論を定数作用素を持つ一般化SDPAに帰着させる。 測定された不変性:相対変動は10^{-6}以下である。摂動界はキャッシュ推論を定量化するが保証はしない。組み立てられた代理指標は復号マージンを捉え損ねる。 条件付きの三段階機構(ロータリーツワール、集中、行写像縮約)が、公開済みチェックポイント上で測定される。大域グラムの下でのブロック単位の訓練とスコアリングは、明示的なターゲット露出を伴って述べられる。試験したサンプルでは、ブロック方式と逐次方式のスコアリングは同一の解答を選択し、公開されているTruthfulQAの確率質量指標に関して項目あたり5×10^{-5}以内で一致する。自己組織化臨界性は、内在的秩序パラメータを持つ現象論的枠組みとして導入され、未解決の主張は反証可能な予想となる。選択された証明コアはLean 4において機械検証されている。
English
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G_{LM}, built from a positive tensor A_{LM} by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at G_{LM}=I; A_{LM} and A_P are strictly entrywise positive, with Perron-Frobenius structure on A_{LM}; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of 10^{-6} and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5times 10^{-5} per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.