Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu

arXiv:2609.35767 · 2026-09-29 공개 · arXiv · PDF

reinforcement-learning image-generation multimodal-reasoning geneval t2i-compbench wise-benchmark unified-models reflection-trajectories

Abstract

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

한국어 요약

한 줄 요약

UMM-Reflection은 단일 통합 모델 내에서 반복적 이미지 수정을 학습하는 강화학습 기반 반사 학습 프레임워크로, GenEval에서 SFT 대비 12.05점 개선을 달성한다.

핵심 기여도

핵심 아이디어

통합 다중모달 모델은 이미지 생성과 분석을 동시에 수행할 수 있으므로, 자체 생성물을 점검하고 수정하는 반사-수정 루프를 구현할 수 있다. 그러나 수정이 도움이 되는지 여부는 렌더링 후에야 확인되므로, 반사 텍스트와 이미지 생성을 전체 루프에서 공동 학습해야 한다. 기존의 단일 렌더링 최적화나 단일 헤드 강화학습은 대부분의 성능 향상을 활용하지 못한다. UMM-Reflection은 하나의 초기 이미지에서 시작하는 여러 경로를 비교하고, 하나의 경로당 반사 토큰과 흐름 기반 수정을 동시에 업데이트함으로써, 라운드별 크레딧 할당의 조합 폭발 문제를 피한다. 이는 단일 라운드 편집이나 외부 평가자 사용 모델과 구별된다.

기술적 접근법

주요 결과

의의 및 한계

UMM-Reflection은 통합 모델 내 반사-수정 루프를 전체적으로 학습하는 첫 사례로, 기존 단일 렌더링 최적화나 외부 평가자 사용 모델과 차별화된다. 특히, 기존 SFT가 이미 생성 가능한 수정 경로 중 성공률 높은 것만 선택함으로써, 새로운 능력이 아닌 기존 능력의 신뢰성 향상을 달성한다는 점에서 학술적 의의가 있다. 그러나 반사 루프 내에서의 복잡한 상호작용을 완전히 모델링하지 못하는 한계가 있으며, 더 긴 반사-수정 루프에 대한 실험은 부재하다.

실용적 활용

UMM-Reflection은 텍스트-이미지 생성 모델의 자가 수정 능력을 향상시켜, 디자인, 콘텐츠 생성, 시각적 분석 등 다양한 산업 분야에서 실용적 활용이 가능하다. 특히, 생성된 이미지의 오류를 자동으로 탐지하고 수정하는 시스템에 적합하며, 사용자 피드백 없이도 반복적인 개선을 수행할 수 있다.