ChatPaper.aiChatPaper

OmniOpt:現代優化器的分類學、幾何學與基準測試

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

July 4, 2026
作者: Siyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang, Xin Jin, Xinglong Xu, Jingxuan Wei, Shengye Pang, Jintao Che, Xuanhe Zhou, Conghui He, Cheng Tan
cs.AI

摘要

大規模模型訓練的優化器選擇已成為一項系統層級的設計決策,受到計算資源、記憶體、調校預算及任務多樣性的共同約束,然而現有上百種方法仍顯得零散混亂。為此,我們提出 OmniOpt,為研究社群提供一份統整的優化器調查與基準食譜。OmniOpt 奠基於四個相互耦合的組件:第一,我們將每一種優化器更新視為透過五階段元管線進行的結構化轉換,並顯示多數方法僅涉及其中一至兩個階段;第二,我們利用範數約束的線性最小化預測器(LMO)來統一不同優化器;第三,這兩個觀點構成一個雙維度分類體系,一維將每種方法歸入機制家族,另一維記錄其欲改善的可量化訓練目標;第四,也是本文的核心,我們將完整的分類體系實例化為一個統一的跨領域基準,涵蓋代表性優化器、模型規模以及從語言模型預訓練到影像分類等訓練場景,系統性地分析各方法家族在多種效果目標上的表現,並揭示其取捨。OmniOpt 因此為研究社群提供了一個在明確機制與目標假設下選擇優化器的操作座標系,並為優化器社群的未來發展指引方向。
English
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remains fragmented. We therefore present OmniOpt, a unified survey and benchmark cookbook of optimizers for the research community. OmniOpt rests on four coupled components. First, we treat every optimizer update as a structured transformation through a five-stage meta-pipeline, and show that most methods engage only one or two of these stages. Second, we use norm-constrained linear minimization oracles (LMOs) to unify different optimizers. Third, these two views ground a dual-dimension taxonomy, one dimension assigning each method to a mechanism family and the other recording the measurable training objectives it aims to improve. Fourth, and at the core of this paper, we instantiate the full taxonomy in a unified cross-domain benchmark spanning representative optimizers, model scales, and training regimes from language model pretraining to image classification, systematically analyzing each method family across multiple effect objectives and laying out their trade-offs. OmniOpt thus supplies the research community with an operational coordinate system for selecting optimizers under explicit mechanism and objective assumptions, and charts a direction for the future development of the optimizer community.