SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie

arXiv:2607.21553 · 2026-07-24 공개 · arXiv · PDF

video-generation diffusion-transformer linear-attention long-sequence vbench high-resolution-video hybrid-attention attnres

Abstract

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

한국어 요약

한 줄 요약

SANA-Video 2.0은 5B 및 14B 규모의 하이브리드 어텐션 비디오 생성 모델로, 720p 영상 생성 시 3.2× 가속된 DiT 성능과 84.30의 VBench 점수를 달성한다.

핵심 기여도

핵심 아이디어

SANA-Video 2.0은 기존 full-softmax 어텐션의 O(N²) 비용 문제를 해결하기 위해 linear attention의 O(N) 확장성을 활용하면서도, 정확한 토큰 상호작용을 복원하기 위해 주기적으로 softmax anchor를 삽입하는 하이브리드 구조를 제안한다. 이는 3:1의 비율로 softmax anchor를 배치하여, linear attention의 효율성과 softmax의 표현력을 결합하는 방식이다. 또한, AttnRes를 통해 anchor에서 생성된 정보를 후속 레이어로 전달함으로써, deep layer의 표현력이 약 12% 향상된다는 점이 핵심이다. 이는 기존 linear attention이 토큰 간 상호작용을 제한적으로 표현한다는 한계를 극복하기 위한 전략이다.

기술적 접근법

주요 결과

의의 및 한계

SANA-Video 2.0은 full-softmax 기반 모델과 유사한 생성 품질을 유지하면서도, O(N) 확장성을 통해 장시간, 고해상도 영상 생성에 유리한 성능을 보인다. 특히, 720p/60s에서 3.2×, 720p/5s에서 3.58×의 추가 가속을 달성한 Sol-Engine 최적화는 실용적 가치가 크다. 그러나, anchor 비율은 proxy study를 통해 결정되었으며, 가장 긴 영상 결과는 텐서 형태의 프로파일링에 의존한다는 점에서 한계가 있다. 또한, 모델의 최적화는 특정 하드웨어 (H100, B200)에 의존적이며, 다른 플랫폼에서의 성능은 추가 연구가 필요하다.

실용적 활용

SANA-Video 2.0은 단일 GPU에서 720p 영상 생성이 가능하며, 3.2× 가속된 DiT 성능을 통해 실시간 또는 저비용 영상 생성에 적합하다. 특히, Sol-Engine 최적화를 통해 120× 빠른 처리가 가능한 점은 클라우드 기반 영상 생성 서비스나 모바일 장치에서의 활용 가능성을 높인다. 또한, AttnRes와 하이브리드 어텐션 구조는 장시간 영상 생성 및 스트리밍 시스템에 적용 가능하다.