ChatPaper.aiChatPaper

OmniOpt: 현대 최적화 알고리즘의 분류 체계, 기하학 및 벤치마킹

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

July 4, 2026
저자: Siyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang, Xin Jin, Xinglong Xu, Jingxuan Wei, Shengye Pang, Jintao Che, Xuanhe Zhou, Conghui He, Cheng Tan
cs.AI

초록

대규모 모델 학습을 위한 최적화기 선택은 연산, 메모리, 튜닝 예산, 작업 다양성에 의해 공동으로 제약되는 시스템 수준 설계 결정이 되었지만, 백 가지가 넘는 방법들이 산재해 있는 상황은 여전히 분산된 상태로 남아 있다. 이에 우리는 연구 커뮤니티를 위한 최적화기의 통합된 조사 및 벤치마크 안내서인 OmniOpt를 제시한다. OmniOpt는 네 가지 상호 연결된 구성 요소에 기반한다. 첫째, 모든 최적화기 갱신을 5단계 메타 파이프라인을 통한 구조적 변환으로 간주하며, 대부분의 방법이 이러한 단계 중 하나 또는 두 개만 활용한다는 것을 보여준다. 둘째, 노름 제약 선형 최소화 오라클(LMO)을 사용하여 서로 다른 최적화기를 통합한다. 셋째, 이 두 관점은 이중 차원 분류 체계의 근거를 마련하는데, 한 차원은 각 방법을 메커니즘 패밀리에 할당하고 다른 차원은 개선하고자 하는 측정 가능한 학습 목표를 기록한다. 넷째, 이 논문의 핵심으로, 전체 분류 체계를 언어 모델 사전 학습부터 이미지 분류까지 대표적인 최적화기, 모델 규모, 학습 체계를 포괄하는 통합된 교차 도메인 벤치마크로 구현하여, 각 방법 패밀리를 여러 효과 목표에 걸쳐 체계적으로 분석하고 절충점을 제시한다. 따라서 OmniOpt는 연구 커뮤니티에 명시적 메커니즘 및 목표 가정 하에 최적화기를 선택할 수 있는 운영 좌표계를 제공하며, 최적화기 커뮤니티의 미래 발전 방향을 제시한다.
English
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remains fragmented. We therefore present OmniOpt, a unified survey and benchmark cookbook of optimizers for the research community. OmniOpt rests on four coupled components. First, we treat every optimizer update as a structured transformation through a five-stage meta-pipeline, and show that most methods engage only one or two of these stages. Second, we use norm-constrained linear minimization oracles (LMOs) to unify different optimizers. Third, these two views ground a dual-dimension taxonomy, one dimension assigning each method to a mechanism family and the other recording the measurable training objectives it aims to improve. Fourth, and at the core of this paper, we instantiate the full taxonomy in a unified cross-domain benchmark spanning representative optimizers, model scales, and training regimes from language model pretraining to image classification, systematically analyzing each method family across multiple effect objectives and laying out their trade-offs. OmniOpt thus supplies the research community with an operational coordinate system for selecting optimizers under explicit mechanism and objective assumptions, and charts a direction for the future development of the optimizer community.