transformer image-classification mamba linear-attention state-space-model dense-prediction vision forget-gate
Abstract
Mamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba's success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba's success, while the other four designs are less crucial. Based on these findings, we propose a Mamba-Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA.
한국어 요약
한 줄 요약
Mamba는 선형 어텐션 트랜스포머와 유사한 구조를 가지며, 잊기 게이트와 블록 설계가 성능 향상에 핵심 역할을 한다.
핵심 기여도
- Mamba를 선형 어텐션 트랜스포머의 변형으로 재구성하여 6가지 주요 차이점을 제시 (입력 게이트, 잊기 게이트, shortcuts, 어텐션 정규화 생략, 싱글헤드, 수정된 블록 설계).
- 실험적으로 잊기 게이트와 블록 설계가 Mamba 성능에 가장 큰 기여를 함을 밝힘.
- Mamba-Inspired Linear Attention (MILA) 모델을 제안하여, 기존 Mamba 모델보다 이미지 분류 및 고해상도 밀도 예측에서 더 우수한 성능을 달성함.
- MILA는 병렬 계산과 빠른 추론 속도를 유지하면서도 Mamba를 능가함.
핵심 아이디어
Mamba는 기존의 선형 어텐션 트랜스포머와 유사한 수학적 구조를 가지지만, 6가지 주요 차이점을 통해 더 높은 성능을 달성한다. 특히, 잊기 게이트는 입력 순서에 민감하게 반응하며, 이는 시퀀스 모델링에서 중요한 역할을 한다. 그러나 이는 순환 계산을 요구하여 비자율적 비주얼 모델에 적합하지 않을 수 있다. 본 연구는 잊기 게이트의 역할을 위치 인코딩으로 대체할 수 있음을 제시하며, 이를 MILA 모델에 적용하였다. Mamba의 블록 설계는 기존 트랜스포머 블록보다 다양한 연산(선택적 SSM, 깊이 합성곱, 활성화 함수 등)을 통합하여 더 효과적인 구조를 제공한다.
기술적 접근법
- Mamba는 선택적 상태 공간 모델(SSM)을 기반으로 하며, 입력 게이트(Δ_i), 잊기 게이트(Ã_i), shortcuts(𝐷 ⊙ 𝑥_i) 등을 포함.
- 선형 어텐션과 Mamba는 동일한 수식 체계로 통합 가능하며, Mamba는 6가지 구조적 차이를 가짐.
- 잊기 게이트는 0~1 사이의 값으로 이전 히든 상태를 감소시키며, 순서 정보를 유지.
- 블록 설계는 H3와 Gated Attention을 결합한 형태로, 선택적 SSM, 깊이 합성곱, 활성화 함수 등을 포함.
- MILA는 잊기 게이트 대신 적절한 위치 인코딩을 사용하여 병렬 계산 가능성을 유지.
주요 결과
- MILA는 ImageNet-1K 데이터셋에서 다양한 Mamba 모델을 능가하며, 고해상도 밀도 예측에서도 우수한 성능을 보임.
- 잊기 게이트와 블록 설계가 Mamba 성능에 가장 큰 기여를 함.
- MILA는 Mamba보다 빠른 추론 속도와 병렬 계산 가능성을 유지하면서도 더 높은 정확도를 달성함.
- Mamba는 순환 계산을 요구하여 비자율적 비주얼 모델에 한계가 있음.
의의 및 한계
- 본 연구는 Mamba의 성공 요인을 선형 어텐션 관점에서 분석함으로써, 기존 트랜스포머와의 차이를 명확히 밝힘.
- 잊기 게이트는 순차적 데이터에 적합하지만, 비자율적 비주얼 모델에서는 병렬 계산을 저해.
- MILA는 이러한 한계를 극복하고, 비주얼 모델에 더 적합한 형태로 Mamba의 장점을 이식.
- 그러나 MILA의 성능은 특정 데이터셋에만 제한적으로 검증되었으며, 더 넓은 범위의 실험 필요.
실용적 활용
- MILA는 고해상도 이미지 분류 및 밀도 예측(예: 세그멘테이션, 객체 인식)에 적합한 모델로 활용 가능.
- 병렬 계산과 빠른 추론 속도를 요구하는 산업용 비주얼 시스템에 유용.
- 비자율적 비주얼 모델 개발에 있어 Mamba의 장점을 활용할 수 있는 새로운 접근법을 제시.