Pyramidal Flow Matching for Efficient Video Generative Modeling

Yang Jin, Zhicheng Sun, Ning Li, Kun Xu, Hao Jiang, Zhuang Nan, Quzhe Huang, Yang Song, Yadong Mu, Zhouchen Lin

arXiv:2410.05954 · 2026-07-27 공개 · arXiv · PDF

video-generation flow-matching diffusion-transformer high-resolution autoregressive-video efficient-modeling pyramidal-flow temporal-pyramid

Abstract

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io.

한국어 요약

한 줄 요약

파이라미달 플로우 매칭을 통해 768p 24fps 10초 영상을 20.7k A100 GPU 시간 내 생성하는 효율적 비디오 생성 모델 제안.

핵심 기여도

핵심 아이디어

기존 비디오 생성 모델은 고해상도 데이터를 다루기 위해 다단계 캐스케이드 구조를 사용하지만, 이는 단계별 최적화로 인해 지식 공유가 어려우며 유연성이 떨어진다. 본 연구는 이 문제를 해결하기 위해 생성 과정을 공간적·시간적 피라미드 단계로 재해석하여 단일 모델에서 통합 최적화할 수 있는 파이라미달 플로우 매칭 알고리즘을 제안한다. 공간 피라미드는 최종 단계만 고해상도에서 작동하도록 설계되어 초기 단계의 계산 부담을 줄이고, 시간 피라미드는 과거 프레임의 정보를 점진적으로 압축하여 토큰 수를 감소시킨다. 이는 단일 Diffusion Transformer(DiT)를 통해 끝에서 끝까지(end-to-end) 최적화가 가능하며, 별도의 슈퍼해상도 모델 없이도 생성과 해제압축을 동시에 수행할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용