Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Liang Wen, Yu Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Juncheng Liu, Li-Juan Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, Xiangzheng Zhang

arXiv:2503.10460 · 2026-07-27 공개 · arXiv · PDF

model-distillation math-reasoning cross-domain-generalization dpo sota curriculum-sft long-cot public-data

Abstract

This paper introduces Light-R1, an open-source suite for training long reasoning models using reproducible and cost-effective methodology. Given the proprietary nature of data used in the DeepSeek-R1 series, we develop an alternative approach leveraging exclusively public data and models. Our curriculum training progressively increases data difficulty, combined with multi-staged post-training. Our Light-R1-32B model, trained from Qwen2.5-32B-Instruct, outperforms DeepSeek-R1-Distill-Qwen-32B in math reasoning. Experimental results show that this curriculum approach becomes more effective when distinct, diverse datasets are available for different training stages: fine-tuning DeepSeek-R1-Distilled models (pre-tuned by DeepSeek team on proprietary data) with 3,000 challenging examples from our curriculum dataset yielded state-of-the-art 7B and 14B models, while the 32B model, Light-R1-32B-DS performed comparably to QwQ-32B and DeepSeek-R1. Furthermore, we extend our work by applying GRPO on long reasoning models. Our final Light-R1-14B-DS achieves SOTA performance among 14B models in math, with AIME24&25 scores of 74.0 and 60.2 respectively, surpassing many 32B models and DeepSeek-R1-Distill-Llama-70B. Despite math-focused training, Light-R1-14B-DS demonstrates strong cross-domain generalization. Light-R1 represents a significant advancement in making sophisticated reasoning models more accessible and implementable in real-world applications. Our models, training data and code have been made available at https://github.com/Qihoo360/Light-R1.

한국어 요약

한 줄 요약

Light-R1은 공개 데이터와 모델을 활용해 재현 가능한 방식으로 장기 추론 모델을 학습하는 오픈소스 툴킷으로, 14B 모델에서 AIME24 점수 74.0을 달성했다.

핵심 기여도

핵심 아이디어

Light-R1은 **DeepSeek-R1 시리즈의 프로퍼티 데이터 의존성**을 극복하기 위해 **공개 데이터와 모델만을 사용한 재현 가능한 학습 전략**을 제시한다. 핵심은 **커리큘럼 학습**(curriculum training)과 **다단계 포스트 트레이닝**(multi-staged post-training)의 결합이다. 학습 초기에는 쉬운 문제부터 점진적으로 어려운 문제로 이동하며, **SFT-Stage 1 → SFT-Stage 2 → DPO**의 단계별 학습을 통해 모델의 추론 능력을 구축한다. 특히, **3,000개의 과제 데이터셋**은 **DeepSeek-R1-Distill 모델**의 성능을 크게 향상시키며, **Light-R1-7B-DS**가 최고 성능 7B 모델로 등장하게 만든다. 또한, **GRPO**(Gradient-based Reward Policy Optimization)를 장기 추론 모델에 적용한 것은 **14B 모델에서 안정적인 응답 길이 증가와 동시에 성능 향상**을 가능하게 했다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

Light-R1은 **프로퍼티 데이터 없이도 장기 추론 모델을 학습**할 수 있다는 점에서 학술적·실용적 의의가 크다. 특히, **14B 모델에서 32B 수준의 성능**을 달성한 것은 **자원 제약 환경에서의 추론 모델 개발**에 중요한 단서를 제공한다. 또한, **커리큘럼 학습과 GRPO의 결합**은 **장기 추론 모델의 학습 효율성과 확장성**을 증명한다. 그러나, **수학 중심 학습으로 인한 과외 도메인 성능 저하**(예: GPQA)는 한계로 지적된다. 이는 **다양한 도메인의 데이터를 통합한 학습 전략**이 필요함을 시사한다.

실용적 활용

Light-R1은 **자원 제약이 있는 엣지 기기나 실시간 응용**에서 활용 가능한 **고성능 추론 모델 개발**에 기여할 수 있다. 특히, **수학 문제 해결, 알고리즘 설계, 과학 분석** 등에서 활용 가능하며, **교육, 연구, 산업 분석** 분야에서 실용적 가치가 높다. 공개된 코드와 데이터셋은 **저비용으로 장기 추론 모델을 학습하고 개선**하려는 연구자와 엔지니어들에게 유용한 자산이다.