ChatPaper.aiChatPaper

멱법칙 그래프 어텐션: 스케일 내적 어텐션의 정확한 일반화, 추론 시 경험적 붕괴

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

August 10, 2026
저자: Burc Gokden
cs.AI

초록

멱법칙 디코더 표현 기반 대규모 언어 모델(PLDR-LLM)과 그 어텐션인 멱법칙 그래프 어텐션(PLGA)은 스케일링된 점곱 어텐션(SDPA)의 고정 이중선형 형식을 학습된 입력 생성 이중선형 연산자 \(G_{LM}\)으로 대체하며, 이는 양의 텐서 \(A_{LM}\)으로부터 요소별 멱법칙에 의해 구성된다. 아키텍처는 완전히 명세화되었으며 고정된 참조 릴리스에 대해 검증되었다. 주장들은 정리, 조건부 정리, 측정, 또는 추측으로 분류된다. 무조건적으로 성립하는 사실은 다음과 같다. PLGA는 \(G_{LM}=I\)에서 정확히 SDPA를 포함한다. \(A_{LM}\)과 \(A_P\)는 엄격히 요소별 양수이며, \(A_{LM}\)은 페론-프로베니우스 구조를 가진다. DAG 정규화기는 NOTEARS 보행 계수 형태를 가지며, 양수성은 정확한 비순환성을 저해한다. 또한 비공명 조건(표준 회전 주파수로 충족됨) 하에서 교환자 기준은 어떤 연산자가 상대 위치 의존성을 보존하는지를 식별한다. 추론 붕괴 정리: 연역적 출력의 정확한 입력 불변성은 추론을 상수 연산자를 가진 일반화된 SDPA로 붕괴시킨다. 측정된 불변성: \(10^{-6}\) 이하의 상대 변동. 섭동 경계는 캐시된 추론을 정량화하지만 보증하지는 않으며, 구성된 프록시는 디코딩 마진을 포착하지 못한다. 조건부 3단계 메커니즘(회전 트월, 집중, 행 사상 수축)이 공개된 체크포인트에서 측정되었다. 전역 그램 하에서의 블록별 훈련 및 점수화는 명시적 대상 노출과 함께 제시된다. 테스트된 샘플에서 블록 및 순차 점수화는 동일한 답변을 선택하며, 공개된 TruthfulQA 확률 질량 지표에서 항목당 \(5\times 10^{-5}\) 이내로 일치한다. 자기조직화 임계성은 내재적 질서 변수를 가진 현상학적 프레임워크로 도입되며, 미해결 주장들은 반증 가능한 추측이 된다. 선별된 증명 핵심부는 Lean 4에서 기계 검증되었다.
English
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G_{LM}, built from a positive tensor A_{LM} by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at G_{LM}=I; A_{LM} and A_P are strictly entrywise positive, with Perron-Frobenius structure on A_{LM}; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of 10^{-6} and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5times 10^{-5} per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.