语义空间的几何:Transformer架构的连续几何框架
The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
July 19, 2026
作者: Zhihua Liang
cs.AI
摘要
我们提出一个连续几何框架,将Transformer架构的离散代数运算建模为语义纤维丛ε=εM×R^d上的积分微分方程(IDE)。从一个简单的几何公理出发——即令牌序列构成一个配备规范测度格点的离散1维流形——我们将现代Transformer的每个核心组件(RMSNorm、RoPE、Softmax注意力、前馈网络、残差流、SGD、权重衰减)转化为微分几何、测度论和随机微积分的连贯词汇表。该框架产生了跨越熵最优传输(将注意力机制视为薛定谔桥)和非平衡热力学(将SGD视为违反细致平衡的伊藤扩散)的定量预测。我们针对五种架构(Qwen3、LLaMA-3.1、Gemma-3、GPT-2、Mistral),横跨1.24亿至80亿参数规模,进行了六部分实验。实证观测结果与几何预测在定量上一致:机器精度下的ε^{-1/2} Lipschitz标度校准(R²=1.000)、Lie-Trotter算子分裂扭转、对称消融不稳定性(证实拓扑稳定性对偶定律)、RoPE环面上庞加莱回归的O(1/k)热力学抑制、热力学上下文极限相变,以及非平衡稳态参数涡旋——这些结果通过两种优化器(AdamW和纯SGD)验证,排除了动量伪影的干扰。结果表明,通过连续随机微分几何的透镜来分析Transformer,为大语言模型的稳定性极限、上下文边界和优化动力学提供了预测性的描述性词汇。
English
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning 124M to 8B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the ε^{-1/2} Lipschitz scaling calibration at machine precision (R^2 = 1.000), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the calO(1/k) thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.