PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, Zhenguo Li

arXiv:2403.04692 · 2026-07-27 공개 · arXiv · PDF

text-to-image image-generation diffusion-transformer high-fidelity token-compression model-efficiency vision-generation weak-to-strong-training

Abstract

In this paper, we introduce PixArt-\Sigma, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-\Sigma represents a significant advancement over its predecessor, PixArt-\alpha, offering images of markedly higher fidelity and improved alignment with text prompts. A key feature of PixArt-\Sigma is its training efficiency. Leveraging the foundational pre-training of PixArt-\alpha, it evolves from the `weaker' baseline to a `stronger' model via incorporating higher quality data, a process we term"weak-to-strong training". The advancements in PixArt-\Sigma are twofold: (1) High-Quality Training Data: PixArt-\Sigma incorporates superior-quality image data, paired with more precise and detailed image captions. (2) Efficient Token Compression: we propose a novel attention module within the DiT framework that compresses both keys and values, significantly improving efficiency and facilitating ultra-high-resolution image generation. Thanks to these improvements, PixArt-\Sigma achieves superior image quality and user prompt adherence capabilities with significantly smaller model size (0.6B parameters) than existing text-to-image diffusion models, such as SDXL (2.6B parameters) and SD Cascade (5.1B parameters). Moreover, PixArt-\Sigma's capability to generate 4K images supports the creation of high-resolution posters and wallpapers, efficiently bolstering the production of high-quality visual content in industries such as film and gaming.

한국어 요약

한 줄 요약

PixArt-Σ는 4K 해상도 이미지를 생성하는 Diffusion Transformer 모델로, 훈련 효율성과 고해상도 생성 능력을 동시에 달성한 모델이다.

핵심 기여도

핵심 아이디어

PixArt-Σ는 기존 모델(PixArt-α)을 기반으로, **weak-to-strong training** 전략을 통해 훈련 효율성을 극대화한 모델이다. 이는 사전 학습된 PixArt-α를 기반으로, 더 높은 품질의 데이터와 KV 압축 기법을 도입함으로써, **더 작은 모델 크기(0.6B)**로도 **4K 해상도 이미지 생성**이 가능하도록 만든다.

기존의 T2I 모델은 대규모 GPU 자원(예: SD1.5는 6000 A100 GPU days)이 필요하지만, PixArt-Σ는 **PixArt-α의 사전 학습 모델을 재사용**하고, **9%의 GPU days**만으로도 1K 해상도 모델을 훈련할 수 있다. 이는 훈련 비용을 대폭 절감하면서도, **고품질 이미지 생성**과 **정밀한 텍스트-이미지 정렬**을 유지한다.

기술적 접근법

주요 결과

의의 및 한계

PixArt-Σ는 고해상도 이미지 생성 모델의 **훈련 비용과 모델 크기를 대폭 줄이면서도 품질을 유지**한 점에서 혁신적인 기여를 한다. 특히, **weak-to-strong training** 전략은 기존 사전 학습 모델을 기반으로 새로운 데이터와 기술을 효율적으로 통합하는 새로운 방식을 제시한다. 이는 AIGC 분야에서 연구자와 개발자들이 **제한된 자원으로도 높은 성능 모델을 개발**할 수 있도록 도와줄 수 있다.

하지만, PixArt-Σ는 **4K 이미지 생성에 최적화**된 모델로, 다른 용도(예: 동영상 생성, 3D 모델링)에는 적용이 제한적일 수 있다. 또한, **Share-Captioner를 사용한 캡션 데이터의 품질**이 모델 성능에 큰 영향을 미치므로, 데이터 품질 관리는 여전히 중요한 과제이다.

실용적 활용

PixArt-Σ는 영화, 게임, 광고 산업에서 **고해상도 포스터, 벽지, 시각 자료 생성**에 활용 가능하다. 특히, **제한된 GPU 자원으로도 빠른 훈련**이 가능하므로, 중소 규모 기업이나 개인 연구자에게 적합한 도구가 될 수 있다. 또한, **정밀한 텍스트-이미지 정렬** 기능은 디자인 작업에서의 효율성 향상에 기여할 수 있다.