Fast Timing-Conditioned Latent Audio Diffusion

Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley, Jordi Pons

arXiv:2402.04825 · 2026-07-27 공개 · arXiv · PDF

latent-diffusion variational-autoencoder audio-generation efficient-inference stablediffusion text-to-music long-form-audio stereo-audio

Abstract

Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding. Further, most previous works do not tackle that music and sound effects naturally vary in their duration. Our research focuses on the efficient generation of long-form, variable-length stereo music and sounds at 44.1kHz using text prompts with a generative model. Stable Audio is based on latent diffusion, with its latent defined by a fully-convolutional variational autoencoder. It is conditioned on text prompts as well as timing embeddings, allowing for fine control over both the content and length of the generated music and sounds. Stable Audio is capable of rendering stereo signals of up to 95 sec at 44.1kHz in 8 sec on an A100 GPU. Despite its compute efficiency and fast inference, it is one of the best in two public text-to-music and -audio benchmarks and, differently from state-of-the-art models, can generate music with structure and stereo sounds.

한국어 요약

한 줄 요약

Stable Audio는 텍스트와 타이밍 조건을 기반으로 44.1kHz 스테레오 오디오를 빠르게 생성하는 라티언트 디퓨전 모델이다.

핵심 기여도

핵심 아이디어

기존 디퓨전 모델은 고정 길이의 오디오만 생성하거나, 계산 비용이 높아 장시간 오디오 생성에 어려움이 있었다. Stable Audio는 **라티언트 디퓨전**(latent diffusion)을 기반으로, **풀 컨볼루셔널 VAE**(Variational Autoencoder)를 사용해 라티언트 공간에서 효율적으로 오디오를 생성한다. 이 모델은 **텍스트 프롬프트**와 **타이밍 임베딩**(timing embeddings)을 조건으로 받아, 생성 오디오의 **내용과 길이를 정밀하게 제어**할 수 있다. 특히, **변수 길이**(variable-length) 오디오 생성이 가능하며, 이는 기존 모델과 차별화된다. 또한, **스테레오 사운드 이펙트**와 **구조화된 음악** 생성을 지원해 응용 범위를 확장한다.

기술적 접근법

주요 결과

의의 및 한계

Stable Audio는 높은 계산 효율성과 빠른 추론 속도를 유지하면서도, **장시간 스테레오 오디오 생성**에 성공한 점에서 학술적·실용적 의의가 있다. 특히, **구조화된 음악**과 **스테레오 사운드 이펙트** 생성 기능은 음악 생성 및 게임 사운드 디자인 분야에서 활용 가능하다. 그러나, **10초 이하의 짧은 오디오**에 최적화된 기존 평가 지표를 **95초 길이의 오디오에 맞게 수정**해야 하며, **캡션의 일관성** 문제도 존재한다. 또한, **44.1kHz 이상의 고해상도 오디오** 생성 가능성은 명시되지 않았다.

실용적 활용

Stable Audio는 **음악 제작**, **게임 사운드 디자인**, **오디오 콘텐츠 자동 생성** 등에 활용 가능하다. 특히, **구조화된 음악**과 **스테레오 사운드 이펙트** 생성 기능은 **크리에이티브 산업**에서 높은 잠재력을 보인다. A100 GPU 기반의 빠른 추론 속도는 **실시간 응용**에도 유리하다.