Demystify Mamba in Vision: A Linear Attention Perspective

Dongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Jun Song, Shiji Song, Bo Zheng, Gao Huang

arXiv:2405.16605 · 2026-07-27 공개 · arXiv · PDF

transformer image-classification mamba linear-attention state-space-model dense-prediction vision forget-gate

Abstract

Mamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba's success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba's success, while the other four designs are less crucial. Based on these findings, we propose a Mamba-Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA.

한국어 요약

한 줄 요약

Mamba는 선형 어텐션 트랜스포머와 유사한 구조를 가지며, 잊기 게이트와 블록 설계가 성능 향상에 핵심 역할을 한다.

핵심 기여도

핵심 아이디어

Mamba는 기존의 선형 어텐션 트랜스포머와 유사한 수학적 구조를 가지지만, 6가지 주요 차이점을 통해 더 높은 성능을 달성한다. 특히, 잊기 게이트는 입력 순서에 민감하게 반응하며, 이는 시퀀스 모델링에서 중요한 역할을 한다. 그러나 이는 순환 계산을 요구하여 비자율적 비주얼 모델에 적합하지 않을 수 있다. 본 연구는 잊기 게이트의 역할을 위치 인코딩으로 대체할 수 있음을 제시하며, 이를 MILA 모델에 적용하였다. Mamba의 블록 설계는 기존 트랜스포머 블록보다 다양한 연산(선택적 SSM, 깊이 합성곱, 활성화 함수 등)을 통합하여 더 효과적인 구조를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용