ChatPaper.aiChatPaper

손실 함수는 기저를 보지 못하지만, Adam은 본다

The Loss Does Not See the Basis, but Adam Does

August 5, 2026
저자: Devender Singh
cs.AI

초록

인수분해된 모델 \(W = UV^\top\)에 대한 경사 하강법은 저랭크 해를 암묵적으로 선호하는 반면, 동일한 작은 초기화에서 시작하는 Adam은 그렇지 않다. 우리는 이 차이를 손실의 게이지 대칭, 즉 \((U,V) \mapsto (UQ, VQ)\) 변환에 대한 불변성에서 찾는다. 경사 흐름의 저랭크 메커니즘은 해당 최적화기가 게이지 등변일 때만 그 최적화기에 이용 가능하다. 이 조건은 전이에 필요조건이지만 저랭크 복원에는 충분조건이 아니다. 경사 하강법, 모멘텀, '공유 스칼라' Adam, Muon, Shampoo는 이를 만족한다. Adam, RMSProp 및 기타 좌표별 방법은 이를 만족하지 못한다. 구조 정리는 무기억 등변 규칙이 바로 그람 행렬에 의해 결정되는 왼쪽 프리컨디셔너들임을 특징화하고, 전이 정리는 경사 흐름의 경로별 속성을 공통 스칼라 흐름으로 전달한다. 그런 다음 우리는 과소결정 행렬 센싱에서 아홉 가지 업데이트 규칙을 심어진 참값에 대한 복원 오류 기준으로 정렬한다. 좌표별에서 공유 스칼라 프리컨디셔닝에 이르는 단일 매개변수 계열은 편향을 단조롭게 복원하며, 이로써 원인이 이방성임을 특정한다. '스펙트럼 스케줄'은 Muon에 대한 두 가지 상반된 보고 사이의 모순을 해소한다: 동일 비율 업데이트는 저랭크 목표를 정확히 복원하지만 스펙트럼 꼬리가 커질수록 그 우위를 잃는다. 트랜스포머에서 Adam은 첫 단계에서 두 개의 게이지 등가 초기화를 갈라놓는 반면, 등변 최적화기들은 두 초기화를 부동소수점 정밀도 수준에서 동일하게 유지하며, Adam은 최종적으로 헤드별 불변량 \(W_Q^\top W_K\)가 상대 프로베니우스 거리에서 56% 차이가 나는 상태에 도달하는데, 이는 어떤 헤드별 회전으로도 좁힐 수 없는 격차다. 두 개의 초분광 데이터셋에서 훈련 손실을 동일하게 맞췄을 때, 경사 하강법은 가장 낮은 샘플링 밀도와 더 낮은 유효 랭크에서 홀드아웃 오류를 43–44% 줄였다. 따라서 기저 선택은 튜닝 세부 사항이 아니라 최적화기가 어떤 보간 함수를 선택할지에 대한 결정이다.
English
Gradient descent on a factored model W = UV^top is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under (U, V) mapsto (UQ, VQ). Gradient flow's low-rank mechanism is available to an optimizer only if that optimizer is gauge-equivariant, a condition necessary for the transfer but not sufficient for low-rank recovery. Gradient descent, momentum, "shared-scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate-wise methods do not. A structure theorem characterizes the memoryless equivariant rules as exactly the Gram-determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common-scalar flows. We then sort nine update rules on underdetermined matrix sensing by recovery error against the planted ground truth. A one-parameter family from coordinate-wise to shared-scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. A "spectral schedule" reconciles two opposing reports about Muon: equal-rate updates recover exactly low-rank targets but lose their edge as the spectral tail grows. In transformers, Adam separates two gauge-equivalent initializations at the first step, where the equivariant optimizers stay at float precision, and ends with the per-head invariants W_Q^top W_K 56% apart in relative Frobenius distance, a gap no per-head rotation can close. On two hyperspectral datasets at matched training loss, gradient descent cuts held-out error by 43-44% at the lowest sampling density, and at lower effective rank. Basis choice is therefore not a tuning detail but a decision about which interpolant the optimizer selects.