VideoMamba: State Space Model for Efficient Video Understanding

Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, Yu Qiao

arXiv:2403.06977 · 2026-07-27 공개 · arXiv · PDF

self-distillation video-understanding multi-modal mamba state-space-model efficient-modeling high-resolution-video long-term-modeling

Abstract

Addressing the dual challenges of local redundancy and global dependencies in video understanding, this work innovatively adapts the Mamba to the video domain. The proposed VideoMamba overcomes the limitations of existing 3D convolution neural networks and video transformers. Its linear-complexity operator enables efficient long-term modeling, which is crucial for high-resolution long video understanding. Extensive evaluations reveal VideoMamba's four core abilities: (1) Scalability in the visual domain without extensive dataset pretraining, thanks to a novel self-distillation technique; (2) Sensitivity for recognizing short-term actions even with fine-grained motion differences; (3) Superiority in long-term video understanding, showcasing significant advancements over traditional feature-based models; and (4) Compatibility with other modalities, demonstrating robustness in multi-modal contexts. Through these distinct advantages, VideoMamba sets a new benchmark for video understanding, offering a scalable and efficient solution for comprehensive video understanding. All the code and models are available at https://github.com/OpenGVLab/VideoMamba.

한국어 요약

한 줄 요약

VideoMamba는 Mamba 기반 SSM을 활용해 비디오 이해를 효율적으로 수행하는 새로운 모델로, 장단기 동작 인식 및 멀티모달 호환성을 뛰어넘는 성능을 보인다.

핵심 기여도

핵심 아이디어

VideoMamba는 기존 3D CNN 및 비디오 트랜스포머의 한계를 극복하기 위해 Mamba의 선택적 상태공간 모델(SSM)을 비디오 분야에 적용한 모델이다. Mamba는 선형 복잡도를 유지하면서 장기 동적 모델링이 가능하다는 점에서 차별화된다. 이 연구는 1D Mamba 블록을 3D 비디오에 확장한 B-Mamba 블록을 도입하여, 공간적 인식 능력을 강화하였다. 또한, Self-Distillation 기법을 통해 대규모 모델의 과적합 문제를 완화하고, 학습 효율성을 높였다.

기술적 접근법

주요 결과

의의 및 한계

VideoMamba는 비디오 이해 분야에서 높은 확장성과 효율성을 동시에 달성한 첫 SSM 기반 모델로, 특히 고해상도 장기 비디오 처리에 적합하다. 또한, Self-Distillation 기법을 통해 대규모 데이터 없이도 성능을 향상시킬 수 있다는 점에서 실용적 가치가 높다. 그러나 연구자들은 현재 모델의 확장성 검증이 미흡하다고 지적하며, 더 큰 모델(예: VideoMamba-g)이나 멀티모달 통합(예: 오디오, 대형 언어 모델)에 대한 연구가 필요하다고 언급했다.

실용적 활용

VideoMamba는 고해상도 장기 비디오 분석, 실시간 동작 인식, 멀티모달 콘텐츠 처리 등 다양한 산업 분야에서 활용 가능하다. 특히, GPU 메모리 효율성과 빠른 처리 속도로 클라우드 기반 비디오 분석 서비스에 적합하며, 연구 분야에서는 대규모 비디오 데이터셋의 학습 효율성 향상에 기여할 수 있다.