Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu

arXiv:2607.11699 · 2026-07-20 공개 · arXiv · PDF

llm-training instruction-following music-generation multilingual-data music-semantic-tokens stereo-rendering audio-quality-metrics text-to-music

Abstract

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.

한국어 요약

한 줄 요약

Qwen-Music은 텍스트와 레퍼런스 음악 기반으로 고음질 음악을 생성하는 대규모 음악 생성 모델로, Melody-CoT와 Qwen-Music-Render를 통해 창의성과 음향 품질을 동시에 향상시킨다.

핵심 기여도

핵심 아이디어

Qwen-Music은 음악 생성을 세미나틱 컴포지션과 고음질 합성의 두 단계로 분리하여 처리한다. Qwen-Music-Tokenizer는 25Hz 단일 코드북 Music Semantic Token으로 음악을 압축하여 LLM이 처리할 수 있는 형태로 변환한다. Qwen-Music-LLM은 Melody-CoT를 통해 멜로디를 먼저 계획한 후 전체 곡을 생성함으로써 창의성과 구조적 일관성을 동시에 달성한다. 이는 기존의 텍스트-음악 생성 모델이 멜로디를 즉석으로 생성하는 방식과 구별된다. Qwen-Music-Render는 이 토큰을 기반으로 스테레오 웨이브폼을 생성하며, Band-Mode Refiner를 통해 고주파 세부 정보를 보완한다. 이는 디스크리트 토큰이 가지는 음질 저하 문제를 해결하는 핵심 기술이다.

기술적 접근법

주요 결과

의의 및 한계

Qwen-Music은 텍스트-음악 생성과 커버송 생성을 통합한 체계적인 프레임워크를 제시하며, Melody-CoT와 Qwen-Music-Render를 통해 음악의 창의성과 음질을 동시에 향상시켰다. 특히, 다국어 음악 데이터를 기반으로 한 모델의 일반화 능력은 음악 생성 모델의 실용성을 높이는 데 기여한다. 그러나, 레퍼런스 음악 기반 생성 시 일부 스타일의 음악에서는 멜로디 보존률이 낮아지는 한계가 존재하며, 이는 향후 연구 주제로 제시된다.

실용적 활용

Qwen-Music은 음악 제작, 음악 교육, 콘텐츠 제작 분야에서 활용 가능하다. 예를 들어, 작곡가가 텍스트로 음악을 생성하거나, 기존 곡을 새로운 스타일로 재해석할 수 있다. 또한, 음악 스트리밍 플랫폼에서 사용자 맞춤형 음악 생성 서비스로도 활용될 수 있다.