DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu

arXiv:2608.31106 · 2026-09-01 공개 · arXiv · PDF

reinforcement-learning high-resolution audio-video-generation multimodal-feedback joint-training audio-video-data-system native-generation gated-cross-modal-attention

Abstract

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.

한국어 요약

한 줄 요약

DreamX-Creator 1.0은 7B 규모의 생성기와 2K 리파이너를 통해 고해상도 동영상과 오디오를 연동 생성하는 오픈 소스 시스템이다.

핵심 기여도

핵심 아이디어

기존 비디오 생성 모델은 오디오를 별도 단계에서 생성하거나 완전히 생략하는 경우가 많아 시각적 동작과 음향 이벤트 간 상호작용이 제한되었다. DreamX-Creator 1.0은 Gated Cross-Modal Attention을 도입하여 오디오와 비디오 스트림이 독립적으로 처리된 후 후반부에 교차 주의 메커니즘을 통해 정보를 공유하도록 설계되었다. 이 메커니즘은 토큰 및 헤드 단위로 활성화 여부를 조절하는 게이트를 사용하여, 모달별 표현력을 유지하면서도 상호작용을 유도한다. 또한, A2V(오디오→비디오), V2A(비디오→오디오), Joint(양방향) 모드를 동일한 훈련 프레임워크 내에서 처리하여 모델의 유연성과 일관성을 동시에 달성한다.

기술적 접근법

주요 결과

의의 및 한계

DreamX-Creator 1.0은 오디오-비디오 생성 모델의 연구 접근성을 높이고, 고해상도 생성 기술의 발전에 기여할 수 있다. 특히, 7B 규모의 생성기와 1-Step 2K Refiner의 공개는 연구자들이 실험 및 재현을 용이하게 할 수 있도록 지원한다. 그러나, 모델의 훈련 데이터셋과 정확한 평가 지표는 명시되지 않았으며, 일부 기능(예: Audio-Video Reinforcement Learning)은 아직 완전한 실증적 검증이 이루어지지 않았다. 또한, 모델의 성능은 특정 유형의 콘텐츠에 따라 변동할 수 있으며, 이에 대한 분석이 추가적으로 필요하다.

실용적 활용

DreamX-Creator 1.0은 콘텐츠 제작, 게임 개발, VR/AR 산업 등에서 실시간 오디오-비디오 생성이 필요한 상황에 적용 가능하다. 또한, 연구자들이 멀티모달 생성 모델의 기초를 탐구하고, 고해상도 비디오 생성 기술을 개선하는 데 활용될 수 있다.