SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang

arXiv:2608.02023 · 2026-08-04 공개 · arXiv · PDF

zero-shot curriculum-learning audio-generation reward-conditioning multi-speaker multi-audio-modality instruct-task swan-data

Abstract

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.

한국어 요약

한 줄 요약

SwanTale은 자연어와 제로샷 기반의 다발성 음성 및 오디오 생성을 지원하는 통합 모델로, SwanData-Caption 데이터셋과 SwanVAE, Unified MoE 등 다양한 기술을 결합하여 뛰어난 표현력을 달성했다.

핵심 기여도

핵심 아이디어

기존 TTS 시스템은 제로샷 기반의 음성 합성에 집중했으나, 창작 분야에서는 자연어 기반의 음성 및 환경 생성이 필요하다. SwanTale은 자연어 캡션과 제로샷 오디오 레퍼런스를 모두 처리할 수 있는 통합 모델로, SwanVAE와 Unified MoE를 통해 다중 오디오 모달리티를 처리한다. 특히, Engram conditioning과 reward-conditioned quality control을 결합하여 생성 품질을 향상시키며, curriculum learning과 GRPO post-training을 통해 점진적으로 모델 능력을 강화한다. 이는 기존 시스템이 단일 작업에만 최적화된 반면, SwanTale은 두 작업을 동시에 처리할 수 있다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

SwanTale은 자연어와 제로샷 기반의 통합 음성 및 오디오 생성을 가능하게 하며, 애니메이션 더빙, 오디오 드라마, 게임 등 다양한 창작 분야에서 활용 가능하다. 특히, 환경 효과와 스피커 성격을 포함한 복합 오디오 생성을 지원하는 점에서 기존 시스템과 차별화된다. 그러나, SwanData-Caption 데이터셋의 수집 비용이 높고, 일부 특수한 스타일은 여전히 부족할 수 있다. 또한, 모델의 복잡성으로 인해 실시간 응용에는 한계가 있을 수 있다.

실용적 활용

SwanTale은 애니메이션 더빙, 오디오 드라마, 게임, 광고, 쇼트비디오 제작 등에서 창작자들이 자연어로 음성 및 환경을 설계하고, 이후 제로샷 기반으로 재사용할 수 있도록 지원한다. 특히, 복잡한 대사와 오디오 효과를 동시에 생성할 수 있어, 오디오 편집 작업을 대폭 줄이고 창의적 표현을 강화할 수 있다.