MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Ali Hatamizadeh, Jan Kautz

arXiv:2407.08083 · 2026-07-27 공개 · arXiv · PDF

vision-language object-detection image-classification image-net semantic-segmentation mamba-transformer vision-backbone mamba-architecture

Abstract

We propose a novel hybrid Mamba-Transformer backbone, MambaVision, specifically tailored for vision applications. Our core contribution includes redesigning the Mamba formulation to enhance its capability for efficient modeling of visual features. Through a comprehensive ablation study, we demonstrate the feasibility of integrating Vision Transformers (ViT) with Mamba. Our results show that equipping the Mamba architecture with self-attention blocks in the final layers greatly improves its capacity to capture longrange spatial dependencies. Based on these findings, we introduce a family of MambaVision models with a hierarchical architecture to meet various design criteria. For classification on the ImageNet-1K dataset, MambaVision variants achieve state-of-the-art (SOTA) performance in terms of both Top-1 accuracy and throughput. In downstream tasks such as object detection, instance segmentation, and semantic segmentation on MS COCO and ADE20K datasets, MambaVision outperforms comparably sized backbones while demonstrating favorable performance. Code: https://github.com/NVlabs/MambaVision

한국어 요약

한 줄 요약

MambaVision은 시각 작업에 최적화된 하이브리드 Mamba-Transformer 백본으로, ImageNet-1K에서 최신 성능을 달성한다.

핵심 기여도

핵심 아이디어

Mamba는 원래 순차적 데이터 처리에 최적화된 SSM 기반 모델이지만, 시각 작업에서는 전체 수용 필드가 필요하다. 이에 따라, MambaVision은 Mamba 블록과 Transformer 블록을 결합한 하이브리드 구조를 제안한다. 특히, 최종 레이어에 self-attention 블록을 추가함으로써, Mamba가 가진 장거리 의존성 처리 능력을 보완한다. 이는 MambaVision Mixer와 MLP를 포함한 새로운 Mamba 블록 설계를 통해 구현되며, 이는 기존 Mamba의 autoregressive 구조를 개선한 것이다. 실험 결과, self-attention 블록을 최종 레이어에 삽입하는 방식이 가장 효과적이며, 이는 장거리 공간 관계를 포착하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

MambaVision은 Mamba와 Transformer의 장점을 결합한 첫 번째 시각 백본으로, 시각 작업에서의 효율성과 성능을 동시에 달성한다. 특히, self-attention 블록을 최종 레이어에 통합함으로써, Mamba의 autoregressive 한 제한을 극복하고, 전역 맥락을 효과적으로 학습할 수 있다. 그러나 MambaVision은 기존 Transformer나 CNN에 비해 복잡한 구조를 가지므로, 훈련 과정에서 추가적인 컴퓨팅 자원이 필요할 수 있다. 또한, 다양한 시각 작업에서의 일반화 성능은 추가 실험을 통해 검증이 필요하다.

실용적 활용

MambaVision은 고해상도 이미지 처리가 필요한 산업 분야(예: 자율주행, 의료 영상 분석)에서 유용하게 활용될 수 있다. 또한, 실시간 성능이 요구되는 애플리케이션에서 높은 이미지 처리량을 통해 실용적 가치를 제공한다. 연구적으로는, Mamba와 Transformer의 하이브리드 설계 패턴을 다른 시각 작업에 확장하는 데 기초가 될 수 있다.