ChatPaper.aiChatPaper

SLAI T-Rex: Ascend SuperPOD上におけるDeepSeek-V4ファミリーの全パラメータポストトレーニング

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

July 22, 2026
著者: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
cs.AI

要旨

兆パラメータ規模のMoEモデルを対象とした全パラメータ事後学習では、大規模分散学習システムにおいて、深刻なメモリ負荷、非重複通信オーバーヘッド、非効率なカーネル実行など、多くのシステムレベルの課題が生じる。多くの大規模LLM学習システムはGPUベースのクラスタ上に構築されているが、本報告書では、Ascend NPU SuperPOD上でのエンドツーエンドの最適化手法を紹介する。DeepSeek-V4モデルファミリーを対象ワークロードとし、モデルレベル並列化、計算通信オーケストレーション、低レベルカーネル実行にわたる階層的最適化フレームワークを開発した。その結果、オープンソースのベースラインレシピと比較して2.93倍の改善となる34.22%のModel FLOPs Utilization (MFU)を達成し、学習安定性も維持した。この最適化されたインフラストラクチャに基づき、複雑なオペレーションズ・リサーチ(OR)タスク向けのCPTおよびSFTワークフローをさらに構築した。本統合フレームワークをSLAI T-Rexと称する。DeepSeek-V4-Flashを基盤として、収集したドメインリソースとソルバー検証済みの合成最適化文書を組み合わせたOR指向のCPTおよびSFTデータパイプラインを開発した。得られたデータセットは、4つのタスクカテゴリと3つの問題表現にわたる10Kの高品質SFTサンプルを含む。本専門モデルは、評価対象モデルの中で最高の平均Zero-shot Pass@1スコア71.81%を達成し、GPT-5.4-MiniおよびベースのDeepSeek-V4-Flashモデルをそれぞれ3.98パーセントポイントおよび11.27パーセントポイント上回った。全体として、本研究は、Ascendインフラ上での効率的な兆パラメータモデル事後学習から、ソルバー基盤の数理モデリングのためのドメイン特化型Flashモデルに至るフルスタックの経路を実証し、複雑な推論のためのフロンティアモデルシステムを前進させるものである。
English
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.