An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi

arXiv:2608.16887 · 2026-08-18 공개 · arXiv · PDF

diffusion-models text-to-image pre-training generative-modeling pixel-space inference-speedup decoder-architecture noise-schedule

Abstract

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.

한국어 요약

한 줄 요약

픽셀 공간 확산 모델의 대규모 훈련 전략을 실증적으로 분석하고, 래티언트-픽셀 전이 레시피를 제안하여 3.18~4.75배의 추론 가속을 달성했다.

핵심 기여도

핵심 아이디어

기존 연구는 픽셀 공간 확산 모델의 대규모 훈련 전략을 명확히 제시하지 못했으며, 대부분 클래스 조건부 또는 소규모 데이터셋에 제한되었다. 본 연구는 래티언트 공간에서 훈련된 모델을 픽셀 공간으로 전이하는 전략을 제안한다. 이는 래티언트 공간에서 생성적 사전지식을 효율적으로 학습한 후, 후기 훈련 단계에서 픽셀 공간으로 전이함으로써 훈련 효율성을 높인다. 핵심 통찰은, 래티언트 공간에서의 훈련이 픽셀 공간보다 수렴 속도가 빠르기 때문에, 이를 기반으로 픽셀 공간 모델을 훈련하는 것이 효과적이라는 점이다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 픽셀 공간 확산 모델의 대규모 훈련 전략을 실증적으로 분석하여, 래티언트-픽셀 전이 레시피를 제시함으로써 추론 효율성과 생성 품질을 동시에 달성하는 실용적 방법을 제시했다. 특히, 래티언트 공간에서의 훈련 효율성을 활용한 전이 전략은 기존 연구에서 다루지 않았던 핵심 기여이다. 그러나, 픽셀 공간 모델의 훈련 속도가 래티언트 공간 모델보다 느리다는 점은 여전히 해결해야 할 문제이며, 이는 대규모 훈련 환경에서의 에너지 소비나 시간 비용에 영향을 줄 수 있다.

실용적 활용

이 연구는 텍스트-이미지 생성 모델의 실시간 응용, 예를 들어 콘텐츠 생성, 디자인 도구, 게임 개발 등에서 빠른 추론이 요구되는 상황에 적용 가능하다. 또한, 래티언트-픽셀 전이 전략은 다른 생성 모델 아키텍처에도 확장 가능하며, 추론 효율성을 요구하는 산업 분야에서 유용하게 활용될 수 있다.