T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, Hongsheng Li

arXiv:2505.00703 · 2026-08-15 공개 · arXiv · PDF

reinforcement-learning chain-of-thought text-to-image token-level wise-benchmark t2i-compbench janus-pro semantic-level

Abstract

Recent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely unexplored. In this paper, we present T2I-R1, a novel reasoning-enhanced text-to-image generation model, powered by RL with a bi-level CoT reasoning process. Specifically, we identify two levels of CoT that can be utilized to enhance different stages of generation: (1) the semantic-level CoT for high-level planning of the prompt and (2) the token-level CoT for low-level pixel processing during patch-by-patch generation. To better coordinate these two levels of CoT, we introduce BiCoT-GRPO with an ensemble of generation rewards, which seamlessly optimizes both generation CoTs within the same training step. By applying our reasoning strategies to the baseline model, Janus-Pro, we achieve superior performance with 13% improvement on T2I-CompBench and 19% improvement on the WISE benchmark, even surpassing the state-of-the-art model FLUX.1. Code is available at: https://github.com/CaraJ7/T2I-R1

한국어 요약

한 줄 요약

T2I-R1은 강화 학습 기반의 이중 수준 CoT를 도입하여 텍스트-이미지 생성 성능을 13~19% 향상시킨 모델이다.

핵심 기여도

핵심 아이디어

기존 텍스트-이미지 생성 모델은 주로 토큰 수준의 생성 과정만 고려했으나, T2I-R1은 생성 과정을 두 가지 수준으로 구분하여 **이중 수준의 CoT**(Chain-of-Thought)를 도입한다.

첫 번째는 **semantic-level CoT**로, 텍스트 프롬프트를 기반으로 이미지의 전체 구조를 계획하는 고수준 추론이다. 예를 들어, 객체의 위치나 형태를 사전에 정의함으로써 생성 과정을 안정화시킨다. 두 번째는 **token-level CoT**로, 패치 단위로 이미지를 생성하는 과정에서 저수준의 픽셀 처리와 시각적 일관성을 유지하는 역할을 한다.

이 두 수준의 CoT는 **BiCoT-GRPO**라는 강화 학습 알고리즘을 통해 통합적으로 최적화된다. 이는 단일 학습 단계에서 **다양한 생성 보상**(vision experts 기반 reward ensemble)을 사용하여 두 수준의 추론을 동시에 학습하게 한다. 이는 기존의 분리된 학습 방식과 달리, 추론과 생성을 유기적으로 연결할 수 있는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

T2I-R1은 텍스트-이미지 생성 분야에서 추론 기반 생성을 가능하게 한 첫 번째 모델로, **고수준 계획**(semantic-level)과 **저수준 생성**(token-level)을 통합적으로 최적화하는 새로운 패러다임을 제시한다. 기존에는 별도의 모델(예: LLM)을 사용하여 프롬프트 해석을 수행하는 방식이 일반적이었으나, T2I-R1은 **단일 ULM 내에서 두 기능을 통합**하여 계산 비용과 복잡도를 줄였다.

하지만, 본 연구는 **비전 전문가 기반 reward ensemble**을 사용하므로, 이 reward 모델의 편향성이나 일반화 능력은 아직 명시되지 않았다. 또한, **비정상적인 상황**(uncommon scenarios)에서의 성능은 향후 연구가 필요하다는 점이 언급된다.

실용적 활용

T2I-R1은 복잡한 텍스트 프롬프트를 해석하고, 시각적 일관성을 유지하면서 고해상도 이미지를 생성해야 하는 **광고, 콘텐츠 제작, 게임 개발** 등 다양한 산업 분야에 적용 가능하다. 특히, 사용자의 **은근한 의도**(true intention)를 추론하여 더 인간 친화적인 결과를 생성하는 능력은 **사용자 맞춤형 이미지 생성** 시스템 개발에 유용하다.