LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie

arXiv:2609.03796 · 2026-09-04 공개 · arXiv · PDF

vision-language image-generation diffusion-transformer muon-optimizer qwen-image-bench photorealistic-images rmsnorm image-only-pretraining

Abstract

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.

한국어 요약

한 줄 요약

LLaDA-Image는 6B 규모의 DiT와 LLaDA2.0-Mini 기반 VLM을 결합한 오픈소스 이미지 생성 프레임워크로, 220M 샘플 중 98%가 실제 이미지이며 Qwen-Image-Bench에서 53.53 점을 달성했다.

핵심 기여도

핵심 아이디어

LLaDA-Image는 기존 이미지-텍스트 쌍에 의존하지 않고, 먼저 이미지 전용 사전 훈련과 중간 훈련을 통해 강력한 시각 생성 사전을 구축한다. 이는 훈련 초기 단계에서 높은 계산 비용을 요구하는 언어 정렬을 피함으로써, 실제 이미지 데이터를 중심으로 훈련을 진행함으로써 더 높은 사실감을 달성하는 전략이다.

핵심적인 통찰은, 텍스트-이미지 생성과 편집을 하나의 체크포인트 내에서 처리할 수 있는 통합 아키텍처를 구축하는 데 있다. 이는 dLLM 기반 VLM과 DiT를 결합한 모듈 구조를 통해 가능하다. 편집 시에는 VLM을 우회하고 DiT 내부에서 직접 참조 이미지 경로를 활성화하여, 편집 정확도를 높인다.

기술적 접근법

주요 결과

의의 및 한계

LLaDA-Image는 오픈소스 이미지 생성 모델의 한계를 극복하는 데 기여한다. 실제 이미지 위주의 훈련 전략과 통합 아키텍처를 통해 높은 사실감과 편집 정확도를 동시에 달성했으며, TwinFlow 기반의 Turbo 버전은 추론 효율성을 확보했다. 또한, 훈련 코드와 체크포인트를 공개함으로써 연구 재현성을 높였다.

그러나, 현재 모델은 1024² 해상도까지만 훈련되었으며, 2K 해상도 생성은 평가 대상이 아님. 또한, 복잡한 레이아웃이나 다중 텍스트 영역 처리 시 안정성이 떨어질 수 있으며, 특정 주제나 문화적 맥락에 대한 지식 표현도 한계가 있다.

실용적 활용

LLaDA-Image는 고해상도 이미지 생성, 텍스트-이미지 생성, 이미지 편집이 필요한 디자인, 콘텐츠 제작, 게임 개발 등 다양한 산업 분야에 적용 가능하다. 특히, 오픈소스로 제공된 훈련 레시피와 코드는 연구자들이 비용 효율적인 생성 모델을 개발하는 데 유용하게 활용될 수 있다.