VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Puyuan Peng, Po-Yao (Bernie) Huang, Shang-Wen Li, Abdelrahman Mohamed, David Harwath

arXiv:2403.16973 · 2026-07-27 공개 · arXiv · PDF

transformer-decoder zero-shot-tts neural-codec speech-editing real-edit valle xtts-v2 audiobook

Abstract

We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. VoiceCraft employs a Transformer decoder architecture and introduces a token rearrangement procedure that combines causal masking and delayed stacking to enable generation within an existing sequence. On speech editing tasks, VoiceCraft produces edited speech that is nearly indistinguishable from unedited recordings in terms of naturalness, as evaluated by humans; for zero-shot TTS, our model outperforms prior SotA models including VALLE and the popular commercial model XTTS-v2. Crucially, the models are evaluated on challenging and realistic datasets, that consist of diverse accents, speaking styles, recording conditions, and background noise and music, and our model performs consistently well compared to other models and real recordings. In particular, for speech editing evaluation, we introduce a high quality, challenging, and realistic dataset named RealEdit. We encourage readers to listen to the demos at https://jasonppy.github.io/VoiceCraft_web.

한국어 요약

한 줄 요약

VoiceCraft는 실생활 음성 편집 및 제로샷 TTS에서 최고 성능을 보이는 신경 코드 언어 모델이다.

핵심 기여도

핵심 아이디어

VoiceCraft는 기존 음성 시퀀스 내에서 토큰을 인필링하여 편집 및 생성을 수행하는 신경 코드 언어 모델이다. 이 모델은 **인과 마스킹**과 **지연 스택**을 결합한 토큰 재배열 절차를 통해, 양방향 컨텍스트를 유지하면서도 자동회귀적 생성이 가능하도록 설계되었다. 이는 기존의 단방향 마스킹 기반 모델과 달리, 편집된 텍스트와 기존 음성 사이의 일관성을 유지하는 데 기여한다. 특히, **RealEdit** 데이터셋은 다양한 발음, 스타일, 배경 소음 등을 포함하여 모델의 실용성을 평가하는 데 중요한 역할을 한다.

기술적 접근법

주요 결과

의의 및 한계

VoiceCraft는 실생활 음성 편집과 제로샷 TTS에서 높은 자연스러움과 정확도를 제공하며, **RealEdit** 데이터셋은 모델 평가의 현실성을 높이는 데 기여한다. 그러나 모델이 특정 발음이나 스타일에 대한 과적합이 발생할 가능성은 명시되지 않았으며, 대규모 음성 생성 시의 계산 효율성에 대한 평가도 부재하다. 또한, **RealEdit**은 310개 샘플로 비교적 작아, 더 큰 데이터셋에서의 성능 검증이 필요하다.

실용적 활용

VoiceCraft는 오디오북, 인터넷 동영상, 팟캐스트 등 다양한 음성 콘텐츠 편집 및 생성에 활용 가능하다. 특히, 제로샷 TTS 기능은 사전 학습 없이 새로운 음성을 생성할 수 있어, 저비용 음성 생성 서비스나 개인 맞춤형 음성 합성에 적합하다.