优化中的渗流动力学:方差级联与离散标度不变性
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
September 2, 2026
作者: Sai Niranjan Ramachandran, Suvrit Sra
cs.AI
摘要
我们研究随机梯度下降(SGD)的动力学。已知SGD会将深度神经网络引导至对应于更简单子网络的不变集。然而,这种引导如何随时间展开仍鲜为人知。我们通过将随机梯度流(SGF)建模为一种渗流过程来回答这一问题;在该过程中,架构对称性迫使子网络以离散的同步块而非逐个方式合并。这些结构转变在宏观序参量中表现为方差尖峰,与物理相变相呼应。我们进一步证明,这种俘获机制及其相关的标度级联在显式重尾噪声模型下可推广到Adam和AdamW。
English
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.