VACE: All-in-One Video Creation and Editing

Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, Yu Liu

arXiv:2503.07598 · 2026-07-27 공개 · arXiv · PDF

video-generation diffusion-transformer video-editing video-synthesis masked-video-editing reference-to-video context-adapter video-condition-unit

Abstract

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of image content creation. However, due to the intrinsic demands for consistency across both temporal and spatial dynamics, achieving a unified approach for video synthesis remains challenging. We introduce VACE, which enables users to perform Video tasks within an All-in-one framework for Creation and Editing. These tasks include reference-to-video generation, video-to-video editing, and masked video-to-video editing. Specifically, we effectively integrate the requirements of various tasks by organizing video task inputs, such as editing, reference, and masking, into a unified interface referred to as the Video Condition Unit (VCU). Furthermore, by utilizing a Context Adapter structure, we inject different task concepts into the model using formalized representations of temporal and spatial dimensions, allowing it to handle arbitrary video synthesis tasks flexibly. Extensive experiments demonstrate that the unified model of VACE achieves performance on par with task-specific models across various subtasks. Simultaneously, it enables diverse applications through versatile task combinations. Project page: https://ali-vilab.github.io/VACE-Page/.

한국어 요약

한 줄 요약

VACE는 다양한 비디오 생성 및 편집 작업을 하나의 모델로 통합한 All-in-One 프레임워크로, VCU와 Context Adapter를 통해 다중 조건 입력을 유연하게 처리한다.

핵심 기여도

핵심 아이디어

VACE는 비디오 생성과 편집 작업을 하나의 모델에서 처리하기 위해 다중 모달 입력을 통합하는 새로운 접근법을 제시한다. 기존에는 각 작업별 모델이 필요했지만, VACE는 VCU라는 인터페이스를 통해 이미지, 비디오, 마스크, 텍스트 등을 하나의 조건으로 통합하여 다중 작업을 처리한다. 특히, Context Adapter라는 구조를 도입하여 시간-공간 차원의 정보를 형식화하여 모델에 주입함으로써, 다양한 작업 개념을 유연하게 처리할 수 있도록 한다. 이는 기존의 작업별 모델과 비교해도 성능을 유지하면서도 사용성을 대폭 개선한다.

기술적 접근법

주요 결과

의의 및 한계

VACE는 비디오 생성 및 편집 분야에서 사용자 경험을 대폭 개선하는 All-in-One 모델로서, 다양한 작업을 하나의 모델로 처리함으로써 서비스 배포 비용과 사용자 상호작용 복잡도를 줄인다. 특히, VCU와 Context Adapter를 통해 시간-공간 일관성을 유지하면서도 유연한 작업 처리가 가능하다는 점에서 학술적·실용적 가치가 크다. 그러나, 다중 작업을 통합한 모델의 복잡성과 추론 시간 증가가 한계로 작용할 수 있으며, 아직 공개된 다중 작업 기준 데이터셋이 부족한 점도 개선이 필요한 부분이다.

실용적 활용

VACE는 영상 콘텐츠 제작, 광고, 게임, VR/AR 등 다양한 산업에서 사용자 맞춤형 비디오 생성 및 편집을 지원할 수 있다. 특히, 복합적인 조건 입력을 필요로 하는 장면 재구성, 다중 조건 기반 생성, 연속 편집 작업 등에서 유연한 활용이 가능하다.