Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, C. Ruan, Ming Zhang, W. Liang, Wangding Zeng

arXiv:2502.11089 · 2026-07-27 공개 · arXiv · PDF

long-context sparse-attention training-efficiency model-optimization token-compression algorithm-design neural-architecture hardware-aligned

Abstract

Long-context modeling is crucial for next-generation language models, yet the high computational cost of standard attention mechanisms poses significant computational challenges. Sparse attention offers a promising direction for improving efficiency while maintaining model capabilities. We present NSA, a Natively trainable Sparse Attention mechanism that integrates algorithmic innovations with hardware-aligned optimizations to achieve efficient long-context modeling. NSA employs a dynamic hierarchical sparse strategy, combining coarse-grained token compression with fine-grained token selection to preserve both global context awareness and local precision. Our approach advances sparse attention design with two key innovations: (1) We achieve substantial speedups through arithmetic intensity-balanced algorithm design, with implementation optimizations for modern hardware. (2) We enable end-to-end training, reducing pretraining computation without sacrificing model performance. As shown in Figure 1, experiments show the model pretrained with NSA maintains or exceeds Full Attention models across general benchmarks, long-context tasks, and instruction-based reasoning. Meanwhile, NSA achieves substantial speedups over Full Attention on 64k-length sequences across decoding, forward propagation, and backward propagation, validating its efficiency throughout the model lifecycle.

한국어 요약

한 줄 요약

NSA는 하드웨어 최적화와 끝에서 끝까지 학습 가능한 희소 어텐션 구조를 통해 64k 길이의 시퀀스에서 1.6× 이상 가속화된 효율적인 롱컨텍스트 모델링을 실현한다.

핵심 기여도

핵심 아이디어

NSA는 기존 희소 어텐션의 주요 한계, 즉 하드웨어 최적화 부족과 훈련 과정에서의 지원 부족을 해결하기 위해 설계되었다. 기존 희소 어텐션 방법은 이론적 계산 감소를 실제 속도 향상으로 전환하지 못하거나, 훈련 과정에서 희소성 패턴을 효과적으로 활용하지 못하는 문제가 있었다. NSA는 **계층적 토큰 모델링**을 통해 이 문제를 해결한다.

NSA는 **coarse-grained token compression**과 **fine-grained token selection**을 결합한 **dynamic hierarchical sparse strategy**를 채택한다. 이는 토큰 집합을 시간별 블록으로 분할하고, 세 가지 어텐션 경로 — 압축된 토큰, 선택된 토큰, 슬라이딩 윈도우 — 를 통해 계산을 분산한다. 이 구조는 글로벌 컨텍스트 유지와 로컬 정밀도를 동시에 보장한다.

또한, NSA는 **arithmetic intensity-balanced algorithm design**을 통해 하드웨어(예: A100 GPU)의 연산-메모리 비율을 균형 있게 유지함으로써, 실제 속도 향상이 가능하도록 설계되었다. 이는 훈련 및 추론 단계에서 모두 적용된다.

기술적 접근법

주요 결과

의의 및 한계

NSA는 롱컨텍스트 모델링에서 계산 효율성과 모델 성능을 동시에 달성하는 중요한 기술적 진전이다. 특히, **하드웨어 맞춤 최적화**와 **학습-추론 통합 설계**는 기존 희소 어텐션의 주요 한계를 극복하고, 실제 대규모 모델에서 적용 가능성을 높인다.

그러나, NSA는 특정 하드웨어(예: A100 GPU)에 최적화된 구조를 가지므로, 다른 하드웨어 환경에서 동일한 성능을 보장하기 위해서는 추가 최적화가 필요할 수 있다. 또한, 훈련 과정에서 희소성 패턴이 모델 성능에 미치는 영향에 대한 심층 분석은 아직 부족하다.

실용적 활용

NSA는 대규모 언어 모델에서 긴 문서, 코드베이스, 다턴 대화 처리 등이 필요한 산업 및 연구 분야에 적용 가능하다. 특히, **실시간 추론**이 요구되는 시스템에서 계산 효율성과 성능을 동시에 달성할 수 있어, 자동화된 에이전트 시스템, 코드 생성, 복잡한 추론 작업 등에 유용하게 활용될 수 있다.