On-Policy Self-Distillation in Diffusion Models

Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua

arXiv:2608.24646 · 2026-08-26 공개 · arXiv · PDF

diffusion-models image-generation post-training alignment on-policy-self-distillation exponential-moving-average reward-guidance sd-3-5-m

Abstract

Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

한국어 요약

한 줄 요약

DiffusionOPSD는 확산 모델의 포스트 트레이닝을 효율적으로 수행하기 위한 온-포로이 자기-디스틸레이션 프레임워크이다.

핵심 기여도

핵심 아이디어

기존의 확산 모델 학습은 종단 보상만을 기반으로 하여 중간 denoising 예측에 대한 구체적인 지침이 부족하다. 이에 반해 DiffusionOPSD는 이미지 수준의 보상 정보를 각 샘플된 쿼리에서 clean-output 예측에 대한 명시적인 타겟으로 변환한다. 이는 정책이 생성한 anchor 주변에서 보상 기울기로 유계한 양·음의 타겟을 구성함으로써 가능하다. 학습 정책은 이 타겟을 분리된 감독으로 사용하며, 이 과정을 유한 학습 후 EMA(지수 이동 평균) 업데이트로 반복한다. 이는 타겟 생성과 학습 과정을 분리해 분석 가능성을 높인다.

기술적 접근법

주요 결과

의의 및 한계

DiffusionOPSD는 확산 모델의 포스트 트레이닝을 효율적이고 분석 가능한 방식으로 수행할 수 있는 새로운 접근법을 제시한다. 특히, 이미지 수준의 보상 정보를 중간 감독으로 변환함으로써 정책 학습의 안정성과 수렴 속도를 동시에 개선한다. 그러나, 단일 fitting update 이후의 성능 개선이 target-construction gain과 항상 비례하지 않는다는 한계가 존재하며, 이는 추가 연구가 필요하다.

실용적 활용

DiffusionOPSD는 이미지 생성, 편집, 생성 모델의 인간 선호도 정렬 등에서 활용 가능하다. 특히, GPU 자원이 제한된 환경에서 효율적인 포스트 트레이닝이 필요한 산업 및 연구 분야에 유용하다.