Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

Sai Niranjan Ramachandran, Suvrit Sra

arXiv:2609.02373 · 2026-09-06 공개 · arXiv · PDF

neural-networks stochastic-gradient-descent phase-transitions adam-optimizer heavy-tailed-noise percolation-process variance-cascades discrete-scale-invariance

Abstract

We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.

한국어 요약

한 줄 요약

SGD의 동역학을 투과 현상(percolation)으로 모델링하여, 신경망의 구조적 축소와 불연속적 상전이를 설명한다.

핵심 기여도

핵심 아이디어

SGD는 신경망 파라미터 공간을 점차 간단한 구조로 축소시키는 경향이 있다. 이 연구는 이 축소 과정을 **투과 현상(percolation)**으로 모델링함으로써, 구조적 병합이 **단일 연결이 아닌 블록 단위로 일어난다는 점**을 밝혔다. 이는 **Reeb 그래프**를 사용하여 불변 집합 주변의 경로를 추적하고, 이를 투과 과정으로 재정규화함으로써 가능했다. 특히, **architectural symmetry**가 파라미터를 동일한 집합으로 끌어들이고, 이들이 **동시에 병합(block-merge)**되며, 이는 **macroscopic order parameter의 분산 증가(variance spikes)**로 관측된다. 이는 물리학의 상전이와 유사한 현상으로, **Discrete Scale Invariance (DSI)**를 유발한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용