Flow-GRPO: Training Flow Matching Models via Online RL

Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, Wanli Ouyang

arXiv:2505.05470 · 2026-08-15 공개 · arXiv · PDF

flow-matching text-to-image reward-hacking policy-gradient geneval online-rl sampling-efficiency ode-to-sde

Abstract

We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation.

한국어 요약

한 줄 요약

Flow-GRPO는 온라인 강화학습을 플로우 매칭 모델에 통합한 최초의 방법으로, GenEval 정확도를 63%에서 95%까지 향상시킨다.

핵심 기여도

핵심 아이디어

Flow-GRPO는 기존 플로우 매칭 모델의 결정론적 ODE 기반 샘플링이 강화학습의 확률적 탐색과 충돌한다는 문제를 해결하기 위해 ODE-to-SDE 전략을 도입한다. 이는 원래 모델의 주변 분포를 유지하면서 확률적 샘플링을 가능하게 한다. 또한, Denoising Reduction 전략은 훈련 시 가우시안 제거 단계를 줄여 샘플링 효율성을 높이되, 추론 시에는 전체 단계를 유지함으로써 성능 저하 없이 훈련 속도를 개선한다. 이 두 전략은 강화학습 기반 최적화를 플로우 모델에 적용할 수 있는 기반을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

Flow-GRPO는 강화학습을 플로우 매칭 모델에 적용하는 최초의 시도로, 텍스트-이미지 생성에서 정확도와 샘플링 효율성을 동시에 향상시키는 기반을 제공한다. 특히, KL 제약을 통해 보상 최대화와 이미지 품질 유지 간의 균형을 맞춘 점이 학술적·실용적으로 의미가 크다. 그러나, 훈련 시간이 KL 제약을 사용할 경우 증가한다는 점은 한계로 작용할 수 있다. 또한, 다양한 보상 유형에 대한 일반화 가능성은 추가 실험을 통해 검증이 필요하다.

실용적 활용

Flow-GRPO는 텍스트-이미지 생성, 시각 텍스트 렌더링, 인간 선호도 정렬 등 생성 모델이 정확성과 다양성을 동시에 요구하는 산업 및 연구 분야에 적용 가능하다. 특히, 복잡한 장면 생성이나 정확한 속성 제어가 필요한 디자인, 콘텐츠 제작, 게임 개발 등에 유용하게 활용될 수 있다.