Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu

arXiv:2607.19064 · 2026-07-22 공개 · arXiv · PDF

diffusion-models image-generation image-editing vae high-resolution rectified-flow-matching mage-flow native-resolution

Abstract

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.

한국어 요약

한 줄 요약

Mage-Flow는 4B 규모의 효율적인 이미지 생성 및 편집 기반 모델로, 고해상도 작업을 낮은 비용으로 가능하게 한다.

핵심 기여도

핵심 아이디어

Mage-Flow는 기존 대규모 이미지 생성 모델의 비용 문제를 해결하기 위해 토크나이저-백본-시스템의 통합 설계를 강조한다. Mage-VAE는 디퓨전 스타일의 단일 스텝 인코딩/디코딩과 앵커-라텐트 정규화를 사용해 높은 재구성 품질을 유지하면서 토크나이징 비용을 대폭 줄였다. NR-MMDiT는 고해상도와 다양한 종횡비를 처리하기 위해 네이티브 해상도 패킹(Native-Resolution Packing)을 도입하여, 기존 버킷 기반 훈련의 제약을 극복했다. 이는 단일 체크포인트가 다양한 해상도에 일반화되도록 한다. 또한, CUDA 커널 퓨전을 통해 메모리 트래픽과 커널 런치 오버헤드를 줄여 훈련 효율성을 향상시켰다.

기술적 접근법

주요 결과

의의 및 한계

Mage-Flow는 대규모 모델의 성능을 유지하면서도 4B 규모로 구현함으로써 연구 및 실용화에 적합한 기반을 제공한다. 특히, 토크나이저-백본-시스템의 통합 설계는 고해상도 작업의 효율성을 극대화한다. 그러나 모델의 크기가 4B로 제한된 점은 복잡한 생성 작업에서는 성능 한계가 있을 수 있다. 또한, Turbo 변형은 단 4단계로 작동하므로 정밀한 편집 작업에서는 제약이 있을 수 있다.

실용적 활용

Mage-Flow는 디자인, UI 프로토타이핑, 과학 도표 제작, 대규모 이미지 편집 플랫폼 등에서 실용적으로 활용될 수 있다. 특히, 낮은 비용과 빠른 처리 속도로 클라우드 기반의 실시간 이미지 생성 및 편집 서비스에 적합하다.