d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning

Siyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya Grover

arXiv:2504.12216 · 2026-08-15 공개 · arXiv · PDF

reinforcement-learning diffusion-models large-language-models reasoning supervised-finetuning policy-gradient non-autoregressive mathematical-benchmarks

Abstract

Recent large language models (LLMs) have demonstrated strong reasoning capabilities that benefits from online reinforcement learning (RL). These capabilities have primarily been demonstrated within the left-to-right autoregressive (AR) generation paradigm. In contrast, non-autoregressive paradigms based on diffusion generate text in a coarse-to-fine manner. Although recent diffusion-based large language models (dLLMs) have achieved competitive language modeling performance compared to their AR counterparts, it remains unclear if dLLMs can also leverage recent advances in LLM reasoning. To this end, we propose d1, a framework to adapt pre-trained masked dLLMs into reasoning models via a combination of supervised finetuning (SFT) and RL. Specifically, we develop and extend techniques to improve reasoning in pretrained dLLMs: (a) we utilize a masked SFT technique to distill knowledge and instill self-improvement behavior directly from existing datasets, and (b) we introduce a novel critic-free, policy-gradient based RL algorithm called diffu-GRPO, the first integration of policy gradient methods to masked dLLMs. Through empirical studies, we investigate the performance of different post-training recipes on multiple mathematical and planning benchmarks. We find that d1 yields the best performance and significantly improves performance of a state-of-the-art dLLM. Our code is released at https://dllm-reasoning.github.io/.

한국어 요약

한 줄 요약

d1은 마스킹된 확산 대형 언어 모델(dLLMs)에 강화학습(RL)을 통합하여 추론 성능을 향상시키는 2단계 프레임워크로, diffu-GRPO라는 새로운 정책 경사 알고리즘을 제안한다.

핵심 기여도

핵심 아이디어

기존의 자동회귀(AR) 모델에서 강화학습(RL)이 추론 성능 향상에 효과적이었지만, 비자동회귀 확산 기반 dLLM에서는 적용이 어려웠다. 이는 dLLM이 반복적인 디노이징 과정을 통해 토큰을 생성하기 때문에, 기존의 PPO나 GRPO와 같은 정책 기반 RL 알고리즘이 적용되지 않았기 때문이다. 이를 해결하기 위해, d1은 SFT와 diffu-GRPO를 결합한 2단계 훈련 프레임워크를 제안한다. SFT 단계에서는 마스킹된 데이터셋을 활용해 추론 추적을 학습시키고, diffu-GRPO 단계에서는 정책 경사 기반의 RL 알고리즘을 도입하여, 마스킹된 dLLM의 정책 로그 확률을 효율적으로 추정한다. 이는 랜덤 마스킹을 통해 정책 최적화의 정규화 효과를 얻고, 배치당 경사도 업데이트 수를 늘려 훈련 시간을 줄인다.

기술적 접근법

주요 결과

의의 및 한계

d1은 마스킹된 dLLM에 RL을 적용하는 새로운 가능성을 제시하며, 기존 AR 모델에서의 추론 성능 향상 전략을 비자동회귀 모델로 확장했다는 점에서 학술적 의의가 있다. 특히, diffu-GRPO는 정책 경사 기반 RL을 dLLM에 최초로 적용한 알고리즘으로, 훈련 효율성과 안정성을 동시에 고려한 설계가 독창적이다. 그러나, LLaDA-8B-Instruct는 4096 토큰 길이로 사전 훈련되었음에도 불구하고, RL 훈련 시 더 긴 토큰 길이가 필요하지만, 생성 속도가 느려 확장이 어렵다는 한계가 있다. 또한, 일부 태스크(예: Sudoku)에서는 토큰 길이 증가에 따른 성능 감소가 관찰되어, 모델의 태스크별 적응력에 대한 추가 연구가 필요하다.

실용적 활용

d1은 수학 문제 해결, 계획 수립, 코드 생성 등 추론이 필요한 다양한 AI 응용 분야에 적용 가능하다. 특히, 비자동회귀 모델의 효율성과 RL 기반 추론 향상 기법을 결합한 이 프레임워크는 대규모 언어 모델의 실시간 추론 성능 향상에 기여할 수 있다. 연구자들이 dLLM을 기반으로 한 새로운 추론 모델을 개발하거나, 기존 모델의 추론 능력을 강화학습으로 향상시키려는 경우에 유용하게 활용될 수 있다.