ChatPaper.aiChatPaper

語義空間的幾何學:Transformer 架構的連續幾何框架

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

July 19, 2026
作者: Zhihua Liang
cs.AI

摘要

我們提出了一個連續幾何框架,將Transformer架構的離散代數運算建模為位於語意纖維叢calE = calM × R^d上的一個積分-微分方程式(IDE)。從一個簡單的幾何公理出發——即token序列構成一個配備了典型測度格點的離散1-流形——我們將現代Transformer的每個核心組件(RMSNorm、RoPE、Softmax Attention、FFN、殘差流、SGD、權重衰減)都翻譯成微分幾何、測度論與隨機微積分的統一語彙。由此產生的框架得出了跨越熵最優傳輸(Attention作為薛丁格橋)與非平衡熱力學(SGD作為破壞細緻平衡的伊藤擴散)的量化預測。我們在五種架構(Qwen3、LLaMA-3.1、Gemma-3、GPT-2、Mistral)上進行了六部分實驗,參數規模從1.24億到80億不等。實證觀測量與幾何預測定量一致:機器精度下的ε^{-1/2} Lipschitz縮放校準(R^2 = 1.000)、Lie–Trotter算子分裂扭轉、對稱消融不穩定性(驗證了拓撲穩定性的雙重定律)、RoPE環面上的Poincaré回歸的calO(1/k)熱力學抑制、熱力學上下文極限相變,以及非平衡穩態參數渦流——以上結果通過兩種優化器(AdamW和純SGD)進行了驗證,以排除動量假象。結果表明,透過連續隨機微分幾何的視角分析Transformer,能為大型語言模型的穩定性極限、上下文邊界與優化動態提供一個具預測性的描述性語彙。
English
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning 124M to 8B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the ε^{-1/2} Lipschitz scaling calibration at machine precision (R^2 = 1.000), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the calO(1/k) thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.