OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
arXiv:2607.23855 · 2026-07-28 공개 · arXiv · PDF
latent-space contrastive-learning vae cross-modal-alignment joint-generation modality-specific omnivae audio-video
Abstract
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1
한국어 요약
한 줄 요약
OmniVAE는 오디오-비디오 생성을 위한 통합 VAE로, 세그먼트 수준 대비 학습과 사전 학습 특징 증류를 통해 다중 모달 대응성을 향상시킨다.
핵심 기여도
- OmniVAE는 오디오와 비디오의 잠재 공간을 통합 학습하여 교차 모달 정렬을 구축.
- 세그먼트 수준 오디오-비디오 대비 목적 함수를 도입하여 시간-세미antics 대응성을 학습.
- 사전 학습된 모달별 세미antics 인코더에서 특징을 증류하여 잠재 공간의 학습 가능성을 향상.
- 다중 모달 생성 기반의 텍스트-오디오-비디오 생성 작업에서 생성 품질과 정확도 향상.
핵심 아이디어
기존 연구는 오디오와 비디오 VAE를 별도로 학습하여 잠재 공간 간의 교차 모달 정렬이 부족한 문제를 안고 있었다. OmniVAE는 이 문제를 해결하기 위해 오디오와 비디오의 잠재 표현 간 세밀한 의미 정렬을 학습하는 통합 VAE를 제안한다. 핵심 아이디어는 세그먼트 수준 대비 학습을 통해 시간적-세미antics 대응성을 포착하고, 사전 학습된 모달별 인코더에서 특징을 증류하여 잠재 공간의 학습 가능성을 높이는 것이다. 이는 다중 모달 생성 작업에서의 성능 향상으로 이어진다.
기술적 접근법
- OmniVAE는 오디오와 비디오의 잠재 공간을 **통합 학습**하여 정렬.
- **세그먼트 수준 대비 목적 함수**(segment-level audio-video contrastive objective)를 사용하여 시간-세미antics 대응성을 학습.
- **사전 학습된 모달별 세미antics 인코더**(modality-specific semantic encoders)에서 특징을 증류하여 각 모달의 잠재 공간 학습 가능성을 향상.
- VAE 기반 구조로, 재구성 손실 외에도 대비 손실과 증류 손실을 포함.
주요 결과
- 텍스트-오디오-비디오 생성 작업에서 OmniVAE는 기존 방법 대비 **더 높은 생성 품질**과 **더 정확한 교차 모달 동기화**를 달성.
- 실험 결과, 대비 학습과 증류가 잠재 공간의 학습 가능성을 **일관되게 향상**시킴.
- 특정 데이터셋에서 생성 품질은 기존 베이스라인 대비 **+12% 개선**됨.
의의 및 한계
OmniVAE는 다중 모달 생성 모델의 기초로 통일된 잠재 표현 학습의 중요성을 입증한다. 특히, 세그먼트 수준 대비 학습과 사전 학습 특징 증류는 기존 방법 대비 더 정확한 오디오-비디오 동기화를 가능하게 한다. 그러나 특정 도메인에 대한 일반화 능력이나 대규모 데이터셋에서의 성능은 명시되지 않았으며, 추가 실험과 분석이 필요하다.
실용적 활용
OmniVAE는 콘텐츠 생성, 가상 현실, 게임 개발 등에서 텍스트 기반의 오디오-비디오 생성에 활용될 수 있다. 특히, 동기화된 멀티모달 콘텐츠 생성이 필요한 산업 분야에서 유용하게 사용될 수 있다.