최적화 상태는 어디에 있어야 하는가? 메모리 효율적인 전문가 혼합 학습을 위한 계층적 상태 할당
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
July 21, 2026
저자: Nuemaan Malik
cs.AI
초록
옵티마이저 상태는 mixture-of-experts(MoE) 학습의 메모리 예산에서 가장 큰 단일 항목이다. 6.78B 파라미터 MoE 언어 모델에서 AdamW는 12.6GB의 bfloat16 가중치를 업데이트하기 위해 50.6GB의 first 및 second moment를 유지한다. 본 연구는 SkewAdam을 연구한다. 이 옵티마이저는 MoE의 세 가지 파라미터 집단(밀집 백본, 전문가, 라우터)이 크기와 그래디언트 통계 측면에서 충분히 달라 동일한 상태를 적용해서는 안 된다는 관찰에 기반한다. SkewAdam은 백본(파라미터의 5%)에 대해 float32 모멘텀과 팩터링된 second moment를 유지하고, 전문가(95%)에 대해 팩터링된 second moment만을 유지하며, 라우터(<0.01%)에 대해 정확한 second moment를 유지한다. 결과 상태는 1.29GB로 AdamW의 2.6%에 해당하며, 최대 학습 메모리는 81.4GB에서 31.3GB로 감소하여 40GB 가속기의 예산 내에 들어간다. 동일한 초기화 상태에서 82M 토큰에 걸친 통제된 비교에서 SkewAdam은 검증 perplexity 108.4를 달성하여 AdamW(126.8), Muon(120.2), Lion(393.7)을 앞질렀으며, 라우터 부하 균형을 균일 하한선의 1% 이내로 안정시켰다. 이러한 할당이 해당 perplexity를 얻는 유일한 요인은 아니다. 계층 제거 실험에서는 20배의 상태를 사용해도 동일한 성능을 보였으며, 팩터링된 추정기를 공유하지만 모멘텀을 생략한 Adafactor는 40포인트 뒤처진 수준에서 정체된다. 계층은 정확도 손실 없이 메모리를 절약하며, 정확도는 모멘텀 유지에서 비롯되는데, 이는 균일 옵티마이저도 공유하는 특성이다. 기준 모델들의 학습률을 최적화하면 격차가 좁혀지지만 완전히 사라지지는 않는다. 가장 잘 튜닝된 AdamW는 118.5, 튜닝된 Adafactor는 139.7에 도달한다. 이러한 결과는 옵티마이저 상태가 어디에 위치하는지가 상태의 양만큼이나 중요함을 시사한다.
English
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the three parameter populations of an MoE - the dense backbone, the experts, and the router - differ enough in size and gradient statistics that they should not receive the same state. SkewAdam keeps float32 momentum plus a factored second moment for the backbone (5% of parameters), a factored second moment alone for the experts (95%), and an exact second moment for the router (<0.01%). The resulting state occupies 1.29 GB, 2.6% of AdamW's, and peak training memory falls from 81.4 GB to 31.3 GB, within the budget of a 40 GB accelerator. In a controlled comparison from identical initializations over 82M tokens, SkewAdam reaches validation perplexity 108.4, ahead of AdamW (126.8), Muon (120.2), and Lion (393.7), and settles router load balance to within 1% of its uniform floor. The allocation is not what earns that perplexity: a tier ablation matches it with twenty times the state, and Adafactor, which shares the factored estimator but drops momentum, plateaus 40 points behind. The tiers buy memory at no cost to accuracy; the accuracy comes from keeping momentum, which a uniform optimizer shares too. Sweeping the baselines' learning rates narrows but does not close the gap: the best tuned AdamW reaches 118.5, tuned Adafactor 139.7. Where optimizer state lives, these results suggest, matters at least as much as how much of it there is.