Dion3: 풀스택 직교적 업데이트
Dion3: Full-Stack Orthogonal Updates
August 12, 2026
저자: Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford
cs.AI
초록
Muon 최적화기는 3차 시간(세제곱 시간)이 걸리는 Newton-Schulz 직교화 단계로 인해 상당한 오버헤드 비용을 발생시킵니다. 가중치가 샤딩되면 통신 오버헤드가 이러한 계산 비용을 가중시켜 많은 환경에서 Muon의 이점을 잠식합니다. 우리는 스택의 모든 수준에서 이러한 오버헤드를 겨냥한 Muon의 개정판인 Dion3를 제시합니다. 우리의 Gram Newton-Schulz 알고리즘은 직교화의 FLOP 비용을 줄이고, CuteDSL 커널은 대칭성을 활용하여 이를 가속화하며, 메가배칭 전략은 통신 오버헤드를 줄입니다. 나아가, 우리는 비용을 더욱 절감하는 간단한 업데이트 규칙 변경을 제안합니다. 즉, 각 단계에서 모멘텀 행렬의 일부 행만 선택하여 직교화하는 것입니다. 이 업데이트 규칙은 또 다른 "압축된" Muon 버전인 Dion보다 속도와 성능 모두에서 개선된 결과를 보여줍니다. 전반적으로 Dion3는 Muon이 달성한 손실과 동일하거나 더 나은 손실을 달성하면서도 최적화기 스텝 시간을 최대 6배 단축합니다. Dion3는 Muon의 대체 솔루션으로 dion 패키지(https://github.com/microsoft/dion)를 통해 제공됩니다.
English
The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. We present Dion3, a revision of Muon that targets this overhead at every level of the stack. Our Gram Newton-Schulz algorithm reduces the FLOP cost of orthogonalization, our CuteDSL kernels accelerate it by exploiting symmetry, and our megabatching strategy reduces communication overhead. Moreover, we propose a simple change to the update rule that cuts costs even further: selecting only a fraction of the momentum matrix's rows to orthogonalize at each step. This update rule improves on Dion (another "compressed" version of Muon), in both speed and performance. Overall, Dion3 matches or improves on the loss achieved by Muon but reduces optimizer step time by up to 6x. Dion3 is available via the dion package (https://github.com/microsoft/dion) as a drop-in replacement for Muon.