TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury

arXiv:2607.21606 · 2026-07-28 공개 · arXiv · PDF

diffusion-models compositional-generation reward-alignment t2i-compbench tilt-framework kl-constrained-objective test-time-modification concept-distributions

Abstract

Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.

한국어 요약

한 줄 요약

TILT는 텍스트-이미지 생성 모델에서 복합 조합 프롬프트에 대한 생성을 개선하기 위해 추론 시점에서 모델 고유의 보상 함수를 활용하는 훈련 없이 작동하는 프레임워크이다.

핵심 기여도

핵심 아이디어

TILT는 복합 조합 생성 실패를 조인트 분포와 단일 개념 분포 간의 모드 겹침으로 해석하고, 이를 해결하기 위해 내재적 보상 함수를 정의한다. 이 보상은 모든 개념이 조합적으로 존재하는 샘플을 선호하며, 외부 감독이나 보상 모델 없이 모델 자체에서 유도된다. 이는 확률적 확산 샘플링 과정에서 KL-제약 최적화를 통해 닫힌 형태의 틸트 타겟 분포를 유도하고, 원칙적인 가이드 단계를 제공한다.

TILT는 두 가지 서로 보완적인 알고리즘, TILT-S와 TILT-C를 제안한다. TILT-S는 높은 노이즈 단계에서 계산 효율적이며, TILT-C는 낮은 노이즈 단계에서 개념별 가이드를 정밀하게 제공한다. 이 두 알고리즘을 시간에 따라 결합한 하이브리드 접근법이 기존 방법 대비 더 나은 성능을 보인다. CO3는 TILT의 특별한 경우로 이론적 근거를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

TILT는 복합 조합 생성 문제를 수학적으로 정당화된 보상 정렬 문제로 포뮬레이션하고, 외부 감독 없이 모델 자체의 내재적 정보를 활용하여 샘플링 과정을 가이드한다. 이는 기존 휴리스틱 접근법과 비교해 더 원칙적인 방법을 제시하며, CO3와 같은 기존 방법을 이론적으로 정당화한다.

하지만 TILT는 **Jacobian 기반 가이드**가 추가적인 계산 비용을 유발하며, **모델-유사도 기울기 최적화**가 불안정할 수 있다. 이는 개별 개념 형성에 영향을 줄 수 있으며, 일부 범주에서 BLIP-VQA 점수가 낮게 나타나는 원인으로 작용한다.

실용적 활용

TILT는 텍스트-이미지 생성 외에도 텍스트-오디오 합성, 분자 생성, 다중 속성 편집 등 다양한 조합 생성 문제에 적용 가능하다. 특히, 외부 감독 없이 모델 자체의 내재적 정보를 활용하는 방식은 훈련 비용을 줄이고, 다양한 생성 작업에서 조합 일관성을 강화하는 데 유용하다.