SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
arXiv:2607.20145 · 2026-07-23 공개 · arXiv · PDF
moe-models deepseek-v4 model-parallelism operations-research zero-shot-evaluation ascend-npu full-parameter-post-training sft-data-pipeline
Abstract
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
한국어 요약
한 줄 요약
Ascend NPU 기반으로 DeepSeek-V4 모델의 트릴리언-파라미터 포스트 트레이닝을 최적화한 SLAI T-Rex 프레임워크를 제시한다.
핵심 기여도
- DeepSeek-V4 모델의 트릴리언-파라미터 MoE 포스트 트레이닝을 Ascend SuperPOD에서 34.22% MFU로 성공적으로 수행.
- 기존 오픈소스 베이스라인 대비 2.93배 개선된 MFU 성능 달성.
- 10K 개의 고질량 SFT 샘플로 구성된 OR(Operations Research) 전용 데이터셋 구축.
- DeepSeek-V4-Flash 모델이 Pass@1 점수에서 GPT-5.4-Mini 및 DeepSeek-V4-Flash 베이스 모델 대비 각각 3.98%와 11.27% 상승.
핵심 아이디어
대규모 MoE 모델의 포스트 트레이닝은 메모리 부담, 커뮤니케이션 오버헤드, 커널 효율성 저하 등 시스템 수준의 문제를 유발한다. 기존 GPU 중심의 시스템 대신, Ascend NPU 기반 SuperPOD에서 모델 수준 병렬성, 계산-커뮤니케이션 조율, 저수준 커널 최적화를 통합한 계층적 최적화 프레임워크를 제안한다. 이는 DeepSeek-V4 모델의 트릴리언-파라미터 규모 포스트 트레이닝을 안정적으로 수행할 수 있는 기반을 제공한다. 또한, OR 문제 해결을 위한 CPT 및 SFT 파이프라인을 개발하여 도메인 전용 모델인 DeepSeek-V4-Flash를 구축한다.
기술적 접근법
- **모델 수준 병렬성**: DeepSeek-V4 모델의 트릴리언-파라미터 MoE 구조를 Ascend NPU에 맞게 최적화.
- **계산-커뮤니케이션 조율**: 커뮤니케이션 오버헤드 최소화를 위한 계산과 커뮤니케이션의 오버래핑 전략 적용.
- **저수준 커널 최적화**: NPU 기반 커널 실행 효율성 향상을 위한 하드웨어 맞춤형 최적화 수행.
- **데이터 파이프라인**: OR 도메인 자원과 solver-verified 합성 최적화 문서를 결합한 데이터셋 구축.
주요 결과
- DeepSeek-V4-Flash 모델은 OR 태스크에서 71.81%의 Pass@1 점수 달성.
- GPT-5.4-Mini 대비 +3.98%, DeepSeek-V4-Flash 베이스 모델 대비 +11.27% 개선.
- 10K 개의 SFT 샘플로 구성된 데이터셋은 4개 태스크 범주와 3가지 문제 표현 방식을 포함.
의의 및 한계
SLAI T-Rex는 Ascend 기반 인프라에서 트릴리언-파라미터 모델의 포스트 트레이닝을 가능하게 하는 전반적인 최적화 전략을 제시하며, 복잡한 수학적 추론을 위한 도메인 전용 모델 개발에 기여한다. 그러나, 사용된 데이터셋의 범위와 문제 표현 방식은 제한적이며, 다른 도메인으로의 확장 가능성은 추가 연구가 필요하다.
실용적 활용
SLAI T-Rex는 NPU 기반 클러스터에서 대규모 모델 트레이닝을 요구하는 산업 및 연구 분야, 특히 최적화 문제 해결을 위한 AI 모델 개발에 적용 가능하다.