意味空間の幾何学:Transformerアーキテクチャのための連続幾何学的枠組み
The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
July 19, 2026
著者: Zhihua Liang
cs.AI
要旨
我々は、Transformerアーキテクチャの離散代数演算を、意味的ファイバーバンドル𝒴 = ℳ × ℝ^d上の積分微分方程式(IDE)としてモデル化する連続幾何学的枠組みを提示する。「トークン列が標準的な測度束を備えた離散1次元多様体を形成する」という単一の幾何学的公理から出発し、現代のTransformerのすべての中核的構成要素(RMSNorm、RoPE、Softmax Attention、FFN、Residual Stream、SGD、Weight Decay)を、微分幾何学、測度論、確率解析の統一的語彙へと翻訳する。得られた枠組みは、エントロピー的最適輸送(Attentionをシュレーディンガー橋渡しとして捉える)や非平衡熱力学(SGDを詳細釣り合いを破る伊藤拡散として捉える)に及ぶ定量的予測を生み出す。我々は、124Mから8Bパラメータにわたる5つのアーキテクチャ(Qwen3、LLaMA-3.1、Gemma-3、GPT-2、Mistral)を対象に、6部構成の実験キャンペーンを実施した。実験観測量は幾何学的予測と定量的に一致する。すなわち、機械精度におけるε^{-1/2} Lipschitzスケーリング較正(R² = 1.000)、Lie-Trotter作用素分割トーション、位相的安定性の双対法則を確認する対称アブレーション不安定性、RoPEトーラス上のポアンカレ再帰に対する𝒪(1/k)熱力学的抑制、熱力学的コンテキスト限界相転移、そして非平衡定常状態パラメータ渦であり、これらは2つの最適化手法(AdamWおよびPure SGD)において、モーメンタムアーティファクトを排除した上で検証された。これらの結果は、連続確率微分幾何学のレンズを通してTransformerを解析することが、大規模言語モデルの安定性限界、コンテキスト境界、および最適化ダイナミクスに対する予言的・記述的語彙を提供することを示している。
English
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning 124M to 8B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the ε^{-1/2} Lipschitz scaling calibration at machine precision (R^2 = 1.000), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the calO(1/k) thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.