ChatPaper.aiChatPaper

SLAI T-Rex: Ascend SuperPOD 기반 DeepSeek-V4 계열의 전체 파라미터 사후 훈련

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

July 22, 2026
저자: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
cs.AI

초록

수조 파라미터 규모 MoE 모델의 전체 파라미터 사후 학습은 대규모 분산 훈련에서 심각한 메모리 압박, 중첩되지 않은 통신 오버헤드, 비효율적인 커널 실행 등 상당한 시스템 수준의 문제를 야기합니다. 대부분의 대규모 LLM 훈련 시스템은 GPU 기반 클러스터를 중심으로 구축되지만, 본 보고서는 Ascend NPU SuperPOD에서의 종단간 최적화 사례를 제시합니다. DeepSeek-V4 모델 제품군을 대상 워크로드로 삼아, 모델 수준 병렬화, 연산-통신 조율, 저수준 커널 실행을 포괄하는 계층적 최적화 프레임워크를 개발했습니다. 그 결과, 훈련 안정성을 유지하면서 오픈소스 기준 레시피 대비 2.93배 개선된 34.22%의 모델 FLOPs 활용률(MFU)을 달성했습니다. 이 최적화된 인프라를 기반으로 복잡한 운영 연구(OR) 작업을 위한 CPT 및 SFT 워크플로우를 추가로 구축했습니다. 통합 프레임워크를 SLAI T-Rex라고 명명했습니다. DeepSeek-V4-Flash를 사용하여 수집된 도메인 리소스와 솔버 검증 합성 최적화 문서를 결합한 OR 지향 CPT 및 SFT 데이터 파이프라인을 개발했습니다. 결과 데이터셋은 4개 작업 범주와 3개 문제 표현을 포괄하는 10K 개의 고품질 SFT 샘플로 구성됩니다. 특화된 모델은 평가 대상 모델 중 최고 평균 제로샷 Pass@1 점수인 71.81%를 달성하여, GPT-5.4-Mini와 기본 DeepSeek-V4-Flash 모델을 각각 3.98퍼센트 포인트 및 11.27퍼센트 포인트 능가했습니다. 전반적으로, 본 연구는 Ascend 인프라에서의 효율적인 수조 파라미터 모델 사후 학습부터 솔버 기반 수학적 모델링을 위한 도메인 특화 Flash 모델에 이르는 전체 스택 경로를 제시하며, 복잡한 추론을 위한 프론티어 모델 시스템을 발전시킵니다.
English
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.