ChatPaper.aiChatPaper

SLAI T-Rex:昇騰SuperPOD上DeepSeek-V4系列的全參數後訓練

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

July 22, 2026
作者: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
cs.AI

摘要

全参数后训练万亿参数级MoE模型,给大规模分布式训练带来了系统层面的重大挑战,包括严重的内存压力、非重叠的通信开销以及低效的内核执行。虽然大多数大规模LLM训练系统基于GPU集群构建,但本报告展示了一项基于昇腾NPU SuperPOD的端到端优化实践。以DeepSeek-V4模型系列为目标负载,我们开发了一个层次化优化框架,涵盖模型级并行、计算-通信协同调度以及底层内核执行。由此构建的系统实现了34.22%的模型FLOPS利用率(MFU),相比开源基线方案提升了2.93倍,同时保持了训练稳定性。在此优化基础设施之上,我们进一步为复杂的运筹学(OR)任务建立了持续预训练(CPT)和监督微调(SFT)工作流。我们将这一集成框架称为SLAI T-Rex。利用DeepSeek-V4-Flash,我们开发了面向OR的CPT和SFT数据流水线,将收集的领域资源与求解器验证的合成优化文档相结合。最终数据集包含10K高质量SFT样本,涵盖四种任务类别和三种问题表示形式。该专用模型在评估模型中取得了最高的平均零样本Pass@1分数,达到71.81%,分别超出GPT-5.4-Mini和基础DeepSeek-V4-Flash模型3.98和11.27个百分点。总体而言,本工作展示了从基于昇腾基础设施的高效万亿参数模型后训练,到面向求解器支撑的数学建模的领域专用Flash模型的完整技术路径,推动了面向复杂推理的前沿模型系统发展。
English
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.