ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, Yi Dong

arXiv:2505.24864 · 2026-09-12 공개 · arXiv · PDF

reinforcement-learning large-language-models kl-divergence reference-policy-resetting pass-k-evaluation task-competence long-horizon-rl pro-rl

Abstract

Recent advances in reasoning-centric language models have highlighted reinforcement learning (RL) as a promising method for aligning models with verifiable rewards. However, it remains contentious whether RL truly expands a model's reasoning capabilities or merely amplifies high-reward outputs already latent in the base model's distribution, and whether continually scaling up RL compute reliably leads to improved reasoning performance. In this work, we challenge prevailing assumptions by demonstrating that prolonged RL (ProRL) training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling. We introduce ProRL, a novel training methodology that incorporates KL divergence control, reference policy resetting, and a diverse suite of tasks. Our empirical analysis reveals that RL-trained models consistently outperform base models across a wide range of pass@k evaluations, including scenarios where base models fail entirely regardless of the number of attempts. We further show that reasoning boundary improvements correlates strongly with task competence of base model and training duration, suggesting that RL can explore and populate new regions of solution space over time. These findings offer new insights into the conditions under which RL meaningfully expands reasoning boundaries in language models and establish a foundation for future work on long-horizon RL for reasoning. We release model weights to support further research: https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

한국어 요약

한 줄 요약

ProRL은 KL divergence 제어와 reference policy reset을 통해 기존 모델 분포를 벗어난 새로운 추론 전략을 발견하는 장기 강화학습 방법이다.

핵심 기여도

핵심 아이디어

ProRL은 기존 강화학습이 단기간 훈련 내에서 기존 모델 분포 내의 최적화만 수행한다는 한계를 극복하기 위해, 훈련 시간을 2,000단계 이상으로 확장하고, KL divergence 제어와 reference policy reset을 도입한 새로운 훈련 프레임워크이다. 기존 연구는 대부분 수학 분야에 집중되어 있었고, 훈련 단계가 짧아 모델이 새로운 추론 전략을 탐색할 시간이 부족했다. ProRL은 다양한 STEM, 코드, 논리 퍼즐, 명령 수행 등 다분야 데이터셋을 사용하여 모델이 보다 일반화된 추론 능력을 개발하도록 유도한다. 특히, 기존 모델이 실패하는 상황에서도 ProRL 모델은 100% pass@1 성능을 달성하는 경우가 있으며, 이는 기존 훈련 데이터와의 중복률이 낮은 새로운 추론 경로를 생성했음을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

ProRL은 기존 강화학습이 단기간 훈련 내에서 기존 모델 분포 내의 최적화만 수행한다는 관점을 재정의하며, 장기 훈련을 통해 새로운 추론 경로를 생성할 수 있음을 입증한다. 이는 추가적인 학습 데이터 없이도 모델의 추론 능력을 확장할 수 있음을 시사하며, 추론 중심 모델 개발에 중요한 기초를 제공한다. 그러나 ProRL은 1.5B 파라미터 모델에만 적용되었으며, 더 큰 모델에서의 효과는 아직 검증되지 않았다. 또한, 훈련 시간이 2,000단계 이상으로 길어지므로 컴퓨팅 자원이 제한된 환경에서는 적용이 어려울 수 있다.

실용적 활용

ProRL은 수학, 코드 생성, STEM 추론, 명령 수행 등 다양한 분야에서 고성능 추론 모델을 개발하는 데 활용될 수 있다. 특히, 기존 모델이 실패하는 복잡한 문제 해결에 효과적일 것으로 기대되며, AI 기반 교육, 자동 프로그래밍, 과학 연구 지원 등에 적용 가능하다.