From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

Tianwei Yin, Qiang Zhang, Richard Zhang, William T. Freeman, Frédo Durand, Eli Shechtman, Xun Huang

arXiv:2412.07772 · 2026-07-27 공개 · arXiv · PDF

video-diffusion image-to-video kv-caching autoregressive-transformer vbench-long streaming-generation distribution-matching-distillation video-to-video-translation

Abstract

Current video diffusion models achieve impressive generation quality but struggle in interactive applications due to bidirectional attention dependencies. The generation of a single frame requires the model to process the entire sequence, including the future. We address this limitation by adapting a pretrained bidirectional diffusion transformer to an autoregressive transformer that generates frames on-the-fly. To further reduce latency, we extend distribution matching distillation (DMD) to videos, distilling 50-step diffusion model into a 4-step generator. To enable stable and high-quality distillation, we introduce a student initialization scheme based on teacher’s ODE trajectories, as well as an asymmetric distillation strategy that supervises a causal student model with a bidirectional teacher. This approach effectively mitigates error accumulation in autoregressive generation, allowing long-duration video synthesis despite training on short clips. Our model achieves a total score of 84.27 on the VBench-Long benchmark, surpassing all previous video generation models. It enables fast streaming generation of high-quality videos at 9.4 FPS on a single GPU thanks to KV caching. Our approach also enables streaming video-to-video translation, image-to-video, and dynamic prompting in a zero-shot manner. We release our code and pretrained models.

한국어 요약

한 줄 요약

CausVid는 비동기적 비디오 생성을 위한 자동 회귀 디퓨전 트랜스포머로, 9.4 FPS 속도로 30초 길이의 고해상도 비디오를 생성한다.

핵심 기여도

핵심 아이디어

기존 디퓨전 비디오 모델은 양방향 어텐션으로 인해 단일 프레임 생성 시 전체 시퀀스를 처리해야 하므로, 실시간 및 인터랙티브 애플리케이션에 적합하지 않았다. 이를 해결하기 위해, CausVid는 사전 학습된 양방향 디퓨전 트랜스포머(DiT)를 자동 회귀 트랜스포머로 변환하여, 프레임을 실시간으로 생성할 수 있도록 했다. 또한, **DMD**(Distribution Matching Distillation)를 비디오 데이터로 확장하여 50단계 모델을 4단계로 압축함으로써 지연 시간을 줄였다. 오류 누적 문제를 해결하기 위해, **비대칭 지도 학습**을 도입하여 양방향 교사 모델로 자동 회귀 학습자 모델을 지도하고, **ODE 초기화**를 통해 안정적인 학습을 보장했다.

기술적 접근법

주요 결과

의의 및 한계

CausVid는 자동 회귀 디퓨전 모델로는 최초로 양방향 디퓨전 모델의 생성 품질을 따라잡았으며, 실시간 및 인터랙티브 비디오 생성에 적합한 속도를 달성했다. 특히, **비대칭 지도 학습**과 **ODE 초기화**를 통해 오류 누적 문제를 해결한 점이 학술적 기여로 평가된다. 그러나, **매우 긴 비디오**(30초 이상) 생성 시 품질 저하가 발생하며, **출력 다양성**이 감소하는 한계가 있다. 또한, 현재 VAE 설계로 인해 **5프레임 단위**로 생성해야 하므로, 프레임 단위 VAE를 도입하면 지연 시간을 추가로 줄일 수 있다.

실용적 활용

CausVid는 **로봇 학습**, **게임 렌더링**, **스트리밍 비디오 편집** 등 실시간 및 장기 비디오 생성이 필요한 산업에 적용 가능하다. 또한, **동적 프롬프팅**과 **이미지-비디오 변환** 기능을 통해 인터랙티브 콘텐츠 제작에 활용할 수 있다.