Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan

arXiv:2610.03665 · 2026-10-05 공개 · arXiv · PDF

code-generation llm-training self-distillation math-benchmarks masked-diffusion information-gain diffusion-rl llada-8b

Abstract

Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.

한국어 요약

한 줄 요약

Pivot-SD는 마스킹된 확산 언어 모델(dLM)의 핵심 토큰만 학습해 효율적인 자기-디스틸레이션을 구현한다.

핵심 기여도

핵심 아이디어

기존의 확산 언어 모델 학습은 대부분 최종 텍스트나 전체 디노이징 단계를 기준으로 했지만, 이는 학습에 중요한 **개별 토큰**(commitments)을 무시한다. Pivot-SD는 **디노이징 과정 중 모델의 불확실성이 급격히 줄어드는 시점**(정보 이득이 높은 단계)에 해당하는 토큰, 즉 **피벗**을 학습 대상으로 선정한다. 이는 확산 모델의 고유한 특성인 **마스킹 상태 기록**을 활용한 새로운 접근법이다.

Pivot-SD는 성공적 트레jectory의 피벗에 **교차 엔트로피**(cross-entropy)를, 실패 트레jectory의 피벗에 **타겟드 언유니크**(targeted unlikelihood)를 적용한다. 이는 전체 시퀀스 대신 핵심 토큰만 학습함으로써 **자원 효율성**과 **성능 개선**을 동시에 달성한다.

기술적 접근법

주요 결과

의의 및 한계

Pivot-SD는 확산 언어 모델의 학습 효율성을 획기적으로 향상시키며, **자연어 추론 및 코드 생성** 등 복잡한 작업에서 뛰어난 성능을 보인다. 특히, **온라인 RL과 비교해 40% 미만의 FLOPs 소모**로 자원 효율성을 입증했다. 그러나 **200개 질문만으로 학습**한다는 점에서 **데이터 확장성**에 대한 한계가 있을 수 있다. 또한, **랜덤 선택 또는 모든 마스킹 토큰 학습**은 정확도를 낮추는 것으로 나타나, **피벗 선택의 중요성**이 입증되었다.

실용적 활용

Pivot-SD는 **자원 제약이 있는 환경**(예: 모바일, 임베디드 시스템)에서 복잡한 추론 작업을 수행하는 확산 언어 모델의 학습에 적합하다. **교육, 소프트웨어 개발, 고객 지원** 등 정확한 추론이 필요한 산업 분야에서 활용 가능하다.