In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang

arXiv:2609.32540 · 2026-09-29 공개 · arXiv · PDF

video-generation kv-cache temporal-consistency vbench gpu-parallelism high-fps-video autoregressive-video-diffusion flashforward

Abstract

Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs 1.16--1.69times faster than HiAR and 1.42--2.92times faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.

한국어 요약

한 줄 요약

FlashForward는 자동 회귀 비디오 생성에서 KV 캐시 재사용을 통해 1.42–2.92배 빠른 생성 속도를 달성한 방법이다.

핵심 기여도

핵심 아이디어

기존 자동 회귀 비디오 생성 방식은 각 청크 생성 후 추가 포워드를 통해 KV 캐시를 업데이트했으나, 이는 계산 부담이 컸다. FlashForward는 **denoising 단계에서 이미 생성된 in-flight KV cache**를 바로 재사용함으로써 추가 포워드를 생략한다. 이는 **stage-specific cache**를 다음 청크에 즉시 제공함으로써 **GPU 병렬 처리**를 가능하게 한다. 하지만 이 early availability는 노이즈가 있는 캐시를 사용하므로 외관 및 움직임의 불일치가 발생할 수 있다. 이를 보완하기 위해 **sparse clean anchor latents**를 사전에 생성하여 **two-sided conditioning**을 통해 생성 경로를 안정화한다. 두 메모리는 서로 다른 시간 스케일에서 작동: clean anchor KV는 **coarse, long-range 구조**를 제공하고, stage-matched history는 **fine, recent 변화**를 유지한다.

기술적 접근법

주요 결과

의의 및 한계

FlashForward는 자동 회귀 비디오 생성에서 **캐시 관리 방식의 혁신**을 제시하며, **생성 속도와 품질의 균형**을 유지하는 데 성공했다. 특히, **in-flight KV cache**의 재사용과 **clean anchor KV**의 조합은 기존 방식의 계산 부담을 줄이면서도 품질을 보장한다. 그러나 **early availability**로 인한 노이즈 문제는 여전히 존재하며, 이는 **anchor 간격**이나 **조건 조절**을 통해 부분적으로 해결된다. 또한, **anchor 생성**이 추가적인 계산을 요구하므로, **anchor 밀도** 조절이 성능과 품질의 균형에 중요하다.

실용적 활용

FlashForward는 **실시간 또는 대규모 비디오 생성**이 필요한 산업, 예를 들어 **광고 제작**, **콘텐츠 생성 플랫폼**, **AI 영상 제작 도구** 등에 적용 가능하다. 특히, **고해상도 및 긴 길이의 비디오**를 빠르게 생성해야 하는 상황에서 유용하며, **다중 GPU 환경**에서의 병렬 처리를 통해 확장성이 높다.