ChatPaper.aiChatPaper

의미 공간의 기하학: 트랜스포머 아키텍처를 위한 연속적 기하학적 프레임워크

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

July 19, 2026
저자: Zhihua Liang
cs.AI

초록

우리는 트랜스포머 아키텍처의 이산적 대수 연산들을 의미론적 올다발 \(\mathcal{E} = \mathcal{M} \times \mathbb{R}^d\) 상의 적분-미분 방정식(IDE)으로 모델링하는 연속 기하학적 프레임워크를 제시한다. 단일 기하학적 공리, 즉 토큰 시퀀스가 표준 측도 격자를 갖춘 이산적 1-다양체를 형성한다는 사실로부터 출발하여, 현대 트랜스포머의 모든 핵심 구성 요소(RMSNorm, RoPE, 소프트맥스 어텐션, FFN, 잔차 스트림, SGD, 가중치 감쇠)를 미분기하학, 측도론, 확률미적분학의 일관된 어휘로 변환한다. 결과적으로 얻어진 프레임워크는 엔트로피 최적 수송(어텐션을 슈뢰딩거 다리로 해석)과 비평형 열역학(SGD를 상세 균형을 위반하는 이토 확산으로 해석)에 걸친 정량적 예측을 제공한다. 우리는 1억 2400만 개에서 80억 개의 파라미터에 이르는 다섯 가지 아키텍처(Qwen3, LLaMA-3.1, Gemma-3, GPT-2, Mistral)에 걸쳐 여섯 부분으로 구성된 실험 캠페인을 수행한다. 실험적 관측량은 기하학적 예측과 정량적으로 일치한다: 기계 정밀도 수준의 \(\varepsilon^{-1/2}\) 립시츠 스케일링 보정(\(R^2 = 1.000\)), 리-트로터 연산자 분할 뒤틀림, 위상 안정성의 이중 법칙을 확인하는 대칭적 절제 불안정성, RoPE 원환체 상에서 푸앵카레 재귀의 \(\mathcal{O}(1/k)\) 열역학적 억제, 열역학적 문맥 한계 상전이, 그리고 비평형 정상 상태 파라미터 소용돌이—이는 두 최적화기(AdamW 및 순수 SGD)에 걸쳐 검증되어 모멘텀 인공물을 배제한다. 결과는 연속 확률 미분기하학의 렌즈를 통해 트랜스포머를 분석하는 것이 대규모 언어 모델의 안정성 한계, 문맥 경계, 및 최적화 동역학에 대한 예측적 기술 어휘를 제공함을 보여준다.
English
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning 124M to 8B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the ε^{-1/2} Lipschitz scaling calibration at machine precision (R^2 = 1.000), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the calO(1/k) thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.