MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

arXiv:2610.06801 · 2026-10-09 공개 · arXiv · PDF

video-generation sparse-attention diffusion-transformers latency-reduction token-grouping meta-cached-sparse denoising-speedup attention-probabilities

Abstract

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a 1.80times denoising speedup on Minimax-H3-Base and a 2.32times speedup on 3D asset generation, both with negligible quality loss.

한국어 요약

한 줄 요약

MC-Sparse는 토큰 수준의 정밀한 어텐션 선택과 타일 정렬 그룹핑을 결합한 훈련 없이 가능한 확산 트랜스포머 가속 프레임워크로, 1.80×~2.32×의 가속 성능을 보인다.

핵심 기여도

핵심 아이디어

기존의 블록 수준 스파스 어텐션(BSA)은 토큰 그룹핑, 선택 정확도, 버려진 토큰의 기여도를 포함한 세 가지 요인으로 인해 품질 저하가 발생한다. 이를 해결하기 위해 MC-Sparse는 토큰 수준의 정확한 어텐션 선택과 타일 정렬 쿼리 그룹핑을 결합한다. 토큰 수준 선택은 정확한 어텐션 확률을 기반으로 KV 토큰을 선택하며, 타일 정렬 그룹핑은 GPU 효율성을 유지한다. 또한, 단계 간 정확한 선택과 잔차 보상의 재사용을 통해 계산 비용을 절감한다. 이는 정확도와 효율성을 동시에 달성하는 새로운 접근법이다.

기술적 접근법

주요 결과

의의 및 한계

MC-Sparse는 확산 트랜스포머의 장거리 생성 작업에서 높은 효율성과 정확도를 동시에 달성하는 중요한 기여를 한다. 특히, 훈련 없이 구현 가능하며, 기존 스파스 어텐션의 한계를 분리하고 개선하는 체계적인 접근법을 제시한다. 그러나, 모든 작업 단계에서 정확한 선택을 수행하지 않으며, 재사용 전략이 특정 상황에서 효과가 줄어들 수 있다. 또한, GPU 메모리 제약이 있는 환경에서는 토큰 수준 선택이 어려울 수 있다.

실용적 활용

MC-Sparse는 고해상도 3D 자산 생성, 비디오 생성 등 장거리 시퀀스 생성 작업에 적용 가능하다. 특히, 빠른 추론이 필요한 산업 현장에서 유용하며, 훈련 없이 즉시 사용할 수 있어 연구 및 개발 효율성을 높인다.