VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, Baining Guo

arXiv:2404.10667 · 2026-07-27 공개 · arXiv · PDF

video-generation real-time-generation high-fps avatar-generation audio-driven facial-dynamics face-latent-space talking-face

Abstract

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but also producing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness. The core innovations include a holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos. Through extensive experiments including evaluation on a set of new metrics, we show that our method significantly outperforms previous methods along various dimensions comprehensively. Our method not only delivers high video quality with realistic facial and head dynamics but also supports the online generation of 512x512 videos at up to 40 FPS with negligible starting latency. It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors.

한국어 요약

한 줄 요약

VASA-1은 단일 정적 얼굴 이미지와 음성 입력으로 실시간으로 생생한 말하는 얼굴 영상을 생성하는 모델로, 높은 품질과 40 FPS의 처리 속도를 제공한다.

핵심 기여도

핵심 아이디어

기존 연구는 주로 입술 동기화에 집중했으나, VASA-1은 **전체 얼굴 동작**(입술, 표정, 눈 움직임, 깜빡임 등)과 **머리 움직임**을 통합적으로 모델링하여, 더 자연스러운 말하는 얼굴을 생성한다. 이는 **Diffusion Transformer**를 사용하여, 얼굴 동작을 하나의 **latent variable**로 통합적으로 표현하고 확률 분포를 학습함으로써 가능하다. 기존 방법은 각 요소에 대해 별도의 모델을 사용했지만, VASA-1은 이를 **하나의 모델에서 통합 처리**하여 일관성과 효율성을 높인다.

또한, **3D-aided representation**을 기반으로 얼굴의 **identity, head pose, 3D appearance, holistic facial dynamics** 등을 disentangled하게 분리한 **latent space**를 학습하여, 생성된 얼굴 동작이 입력 이미지의 정체성을 유지하면서도 다양한 감정 표현이 가능하도록 한다.

기술적 접근법

주요 결과

의의 및 한계

VASA-1은 **실시간 음성-얼굴 생성** 분야에서 중요한 진전을 이루었으며, 특히 **입술 동기화**, **표정 다양성**, **머리 움직임 자연성**을 동시에 달성한 점에서 혁신적이다. **CAPP 메트릭** 도입은 음성-머리 움직임 동기화 평가의 기준을 제시하며, **Diffusion Transformer 기반의 실시간 생성**은 기존 기술의 계산 부담 문제를 해결한다.

하지만, **입력 음성의 감정 표현 범위**나 **복잡한 환경**(예: 노이즈, 다중 음성)에서의 성능은 명시되지 않았으며, **더 긴 영상**(예: 수십 분 이상) 생성 시 일관성 유지 여부는 추가 연구가 필요하다.

실용적 활용

VASA-1은 **AI 챗봇**, **가상 교육자**, **의료 상담 AI**, **게임 및 콘텐츠 제작** 등 다양한 분야에서 활용 가능하다. 특히, **실시간 상호작용**이 필요한 환경(예: 온라인 강의, 원격 상담)에서 **생생한 감정 표현**을 가진 가상 인물을 생성할 수 있어, 사용자 경험을 향상시킬 수 있다.