ChatPaper.aiChatPaper

冪律圖注意力:縮放點積注意力的精確推廣及其在推論時的經驗坍縮

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

August 10, 2026
作者: Burc Gokden
cs.AI

摘要

PLDR-LLM(冪律解碼器表示之大型語言模型)及其注意力機制 PLGA(冪律圖注意力)以一個由輸入生成、經學習得到的雙線性算子 G_{LM},取代了縮放點積注意力(SDPA)的固定雙線性形式;該算子係由正張量 A_{LM} 經逐元素冪律建構而成。此架構已被完整定義,並對照固定的參考版本進行驗證;各項主張分別標記為定理、條件定理、測量結果或猜想。無條件成立者包括:PLGA 在 G_{LM}=I 時精確包含 SDPA;A_{LM} 與 A_P 均為嚴格逐元素正,且 A_{LM} 具有佩隆-弗羅貝尼烏斯結構;DAG 正則化器具有 NOTEARS 遊走計數形式,且正性阻礙精確的無環性;在非共振條件下(標準旋轉頻率滿足此條件),交換子判準可識別哪些算子保留相對位置依賴性。推論坍縮定理:演繹輸出的精確輸入不變性會將推論坍縮為具有常數算子的廣義 SDPA。測量的不變性:相對波動在 10^{-6} 及以下;擾動界限雖可量化但無法認證快取推論;所組裝的代理指標遺漏了解碼餘量。一個條件性三階段機制(旋轉扭絞、集中化、列映射收縮)已在公開的檢查點上測量驗證。在全局 Gram 下的分塊訓練與評分有明確的目標揭露聲明;在所測樣本上,分塊評分與序列評分選出相同答案,並在已發表的 TruthfulQA 機率質量指標上於每項 5×10^{-5} 範圍內一致。自組織臨界性作為一種具有內稟序參數的現象學框架引入;開放性主張因而成為可證偽的猜想。部分選定的證明核心已於 Lean 4 中完成機器驗證。
English
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G_{LM}, built from a positive tensor A_{LM} by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at G_{LM}=I; A_{LM} and A_P are strictly entrywise positive, with Perron-Frobenius structure on A_{LM}; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of 10^{-6} and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5times 10^{-5} per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.