Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, Junyang Lin

arXiv:2505.06708 · 2026-07-27 공개 · arXiv · PDF

transformer long-context mixture-of-experts model-scaling attention-sink gated-attention scaled-dot-product-attention sparse-gating

Abstract

Gating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention. Yet, existing literature rarely examines the specific effects of gating. In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants. Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset. Our central finding is that a simple modification-applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)-consistently improves performance. This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties. By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output. Notably, we find this sparse gating mechanism mitigates 'attention sink' and enhances long-context extrapolation performance, and we also release related $\href{https://github.com/qiuzh20/gated_attention}{codes}$ and $\href{https://huggingface.co/QwQZh/gated_attention}{models}$ to facilitate future research.

한국어 요약

한 줄 요약

Gated Attention을 통해 SDPA 출력에 시그모이드 게이트를 적용한 결과, 모델 성능과 훈련 안정성이 향상되며 'attention sink' 현상을 제거할 수 있었다.

핵심 기여도

핵심 아이디어

기존 어텐션 메커니즘은 선형 변환에 제한되어 표현력이 낮고, 토큰 초기부에 과도한 가중치가 집중되는 'attention sink' 문제가 발생한다. 본 연구는 SDPA 출력에 헤드별 시그모이드 게이트를 적용하여 두 가지 핵심 효과를 도출했다. 첫째, 게이트가 SDPA의 저랭크 선형 변환에 비선형성을 도입하여 모델 표현력을 높인다. 둘째, 입력에 따라 희소한 게이트 점수를 적용하여 어텐션 출력의 희소성을 유도함으로써 attention sink를 제거하고, 긴 문맥 처리 능력을 향상시킨다. 이는 기존의 라우팅 기반 게이트와는 별도로, 게이트 자체의 내재적 효과임을 실험적으로 입증했다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 게이트가 단순히 라우팅 기능을 넘어서, 어텐션 표현력과 훈련 안정성에 직접적인 영향을 미친다는 점을 실증적으로 밝혔다. 특히, 희소 게이트가 attention sink를 제거하고 긴 문맥 처리 능력을 향상시킨다는 점은 기존 연구에서 다루어지지 않았던 중요한 통찰이다. 그러나 본 연구는 특정 데이터셋(3.5T 토큰)에서만 실험되었으며, 다른 도메인이나 더 큰 모델에서의 일반화 가능성은 추가 연구가 필요하다. 또한, 게이트가 도입하는 계산 비용은 낮지만, 대규모 배치 처리 환경에서는 미세한 성능 저하가 발생할 수 있다.

실용적 활용

Gated Attention은 대규모 언어 모델의 훈련 안정성과 추론 성능을 동시에 향상시키므로, 특히 긴 문맥 처리가 필요한 챗봇, 문서 요약, 코드 생성 등에 유용하게 활용될 수 있다. 또한, attention sink 문제를 제거함으로써 모델이 초기 토큰에 과도하게 의존하는 현상을 줄일 수 있어, 보다 균형 잡힌 어텐션 분포를 필요로 하는 응용 분야에도 적합하다.