SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo

arXiv:2607.20145 · 2026-07-23 공개 · arXiv · PDF

moe-models deepseek-v4 model-parallelism operations-research zero-shot-evaluation ascend-npu full-parameter-post-training sft-data-pipeline

Abstract

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.

한국어 요약

한 줄 요약

Ascend NPU 기반으로 DeepSeek-V4 모델의 트릴리언-파라미터 포스트 트레이닝을 최적화한 SLAI T-Rex 프레임워크를 제시한다.

핵심 기여도

핵심 아이디어

대규모 MoE 모델의 포스트 트레이닝은 메모리 부담, 커뮤니케이션 오버헤드, 커널 효율성 저하 등 시스템 수준의 문제를 유발한다. 기존 GPU 중심의 시스템 대신, Ascend NPU 기반 SuperPOD에서 모델 수준 병렬성, 계산-커뮤니케이션 조율, 저수준 커널 최적화를 통합한 계층적 최적화 프레임워크를 제안한다. 이는 DeepSeek-V4 모델의 트릴리언-파라미터 규모 포스트 트레이닝을 안정적으로 수행할 수 있는 기반을 제공한다. 또한, OR 문제 해결을 위한 CPT 및 SFT 파이프라인을 개발하여 도메인 전용 모델인 DeepSeek-V4-Flash를 구축한다.

기술적 접근법

주요 결과

의의 및 한계

SLAI T-Rex는 Ascend 기반 인프라에서 트릴리언-파라미터 모델의 포스트 트레이닝을 가능하게 하는 전반적인 최적화 전략을 제시하며, 복잡한 수학적 추론을 위한 도메인 전용 모델 개발에 기여한다. 그러나, 사용된 데이터셋의 범위와 문제 표현 방식은 제한적이며, 다른 도메인으로의 확장 가능성은 추가 연구가 필요하다.

실용적 활용

SLAI T-Rex는 NPU 기반 클러스터에서 대규모 모델 트레이닝을 요구하는 산업 및 연구 분야, 특히 최적화 문제 해결을 위한 AI 모델 개발에 적용 가능하다.