FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du

arXiv:2607.16190 · 2026-07-25 공개 · arXiv · PDF

video-generation diffusion-transformers sparse-attention di-t load-balancing multi-gpu top-p-routing runtime-optimization

Abstract

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present , a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. uses Top-p routing, a Top-k safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, reduces average load imbalance from 1.34 to 1.08 and delivers a 4.41times attention speedup over FlashAttention, while achieving a 2.02--2.11times DiT inference speedup with competitive video quality.

한국어 요약

한 줄 요약

FVAttn은 다중 GPU 환경에서 실행 효율성을 향상시키는 적응형 희소 어텐션 시스템으로, 런타임 로드 밸런싱과 슬랙 인식 희소 증강을 통해 어텐션 속도를 4.41배 개선한다.

핵심 기여도

핵심 아이디어

기존 Top-p 라우팅은 어텐션 헤드별로 유지되는 블록 수가 달라져 다중 GPU 환경에서 로드 불균형을 유발한다. 이는 동기화 단계에서 일부 랭크가 대기하게 되는 **랭크 레벨 스트래글러 문제**를 초래한다. FVAttn은 이 문제를 해결하기 위해 **런타임에 실제 어텐션 패턴을 기반으로 헤드 재할당**하는 RLB를 도입한다. RLB는 P2P 통신을 통해 과부하 헤드를 이동시키며, 각 랭크는 최대 1개의 헤드만 전송/수신하여 20% 이하의 추가 부담만 유발한다. 또한, SASA는 잔여 슬랙을 활용해 **추가 고가치 블록을 삽입**하여, 전체 어텐션 커버리지를 향상시키며 지연 증가 없이 품질을 개선한다. 이는 기존의 히스토리 기반 레이아웃 재설계와 달리, **현재 실행 중인 어텐션 패턴을 기반으로 즉각적이고 경량한 복구**가 가능하다는 핵심 통찰이다.

기술적 접근법

주요 결과

의의 및 한계

FVAttn은 희소 어텐션의 정확도를 유지하면서도 **다중 GPU 환경에서의 실행 효율성**을 극대화하는 기술로, 특히 **짧은 단계 수 비디오 생성**에서 유용하다. 기존 히스토리 기반 레이아웃 재설계와 달리, **현재 실행 패턴을 기반으로 즉각적 복구**를 수행하므로, 동적 희소 라우팅 환경에서 더 효과적이다. 그러나 **GPU 수가 적은 경우 (예: 2-GPU)**, 초기 불균형이 낮아 RLB의 효과가 제한적일 수 있다. 또한, **P2P 통신 비용**이 높은 환경에서는 RLB의 성능 향상이 감소할 수 있다.

실용적 활용

FVAttn은 고해상도 비디오 생성을 위한 Video DiT 모델에서 실시간 추론 성능을 향상시키는 데 유용하다. 특히, 단계 수가 적은 추론 환경 (예: step-distilled 모델)에서 높은 효율성을 발휘하며, 멀티 GPU 서버에서의 배포 및 실행 최적화에 적합하다. 이는 클라