Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips

Man Yao, Jiakui Hu, Tianxiang Hu, Yifan Xu, Zhaokun Zhou, Yonghong Tian, Boxing Xu, Guoqi Li

arXiv:2404.03663 · 2026-07-27 공개 · arXiv · PDF

transformer image-classification image-net spiking-neural-networks vision-tasks spike-driven neuromorphic-chips meta-architecture

Abstract

Neuromorphic computing, which exploits Spiking Neural Networks (SNNs) on neuromorphic chips, is a promising energy-efficient alternative to traditional AI. CNN-based SNNs are the current mainstream of neuromorphic computing. By contrast, no neuromorphic chips are designed especially for Transformer-based SNNs, which have just emerged, and their performance is only on par with CNN-based SNNs, offering no distinct advantage. In this work, we propose a general Transformer-based SNN architecture, termed as ``Meta-SpikeFormer", whose goals are: 1) Lower-power, supports the spike-driven paradigm that there is only sparse addition in the network; 2) Versatility, handles various vision tasks; 3) High-performance, shows overwhelming performance advantages over CNN-based SNNs; 4) Meta-architecture, provides inspiration for future next-generation Transformer-based neuromorphic chip designs. Specifically, we extend the Spike-driven Transformer in \citet{yao2023spike} into a meta architecture, and explore the impact of structure, spike-driven self-attention, and skip connection on its performance. On ImageNet-1K, Meta-SpikeFormer achieves 80.0\% top-1 accuracy (55M), surpassing the current state-of-the-art (SOTA) SNN baselines (66M) by 3.7\%. This is the first direct training SNN backbone that can simultaneously supports classification, detection, and segmentation, obtaining SOTA results in SNNs. Finally, we discuss the inspiration of the meta SNN architecture for neuromorphic chip design. Source code and models are available at \url{https://github.com/BICLab/Spike-Driven-Transformer-V2}.

한국어 요약

한 줄 요약

Meta-SpikeFormer는 ImageNet-1K에서 80.0% 정확도를 달성한 첫 Transformer 기반 SNN이며, CNN 기반 SNN 대비 3.7% 개선된 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 CNN 기반 SNN은 spike-driven 패러다임을 따르지만, Transformer 기반 SNN은 여전히 Multiply-and-Accumulate (MAC) 연산을 사용하여 에너지 효율성이 떨어졌다. 본 연구는 spike-driven self-attention (SDSA)를 도입하여 Transformer 기반 SNN의 에너지 소모를 줄이고, Conv와 Transformer 기반 블록을 결합한 meta architecture를 설계함으로써, CNN 기반 SNN 대비 뛰어난 성능과 범용성을 달성했다. 특히, SDSA 연산은 전체적으로 sparse addition만 수행하며, 이는 기존 Transformer의 softmax, dot-product 연산과 달리 에너지 효율적이다. 또한, skip connection과 블록 구조의 조합을 통해 다양한 비전 작업을 처리하는 데 효과적임을 입증했다.

기술적 접근법

주요 결과

의의 및 한계

Meta-SpikeFormer는 Transformer 기반 SNN의 성능과 범용성을 획기적으로 향상시켜, CNN 기반 SNN의 한계를 극복하였다. 특히, spike-driven self-attention 연산자는 기존 MAC 기반 Transformer와 달리 에너지 효율성을 유지하면서도 뛰어난 성능을 보인다. 이는 차세대 neuromorphic 칩 설계에 중요한 참고 자료가 될 수 있다. 그러나 본 연구는 아직 실제 neuromorphic 칩에서의 구현 가능성에 대한 구체적인 평가가 부족하며, 대규모 SNN의 에너지 소모 및 실시간 처리 능력에 대한 추가 연구가 필요하다. 또한, 다양한 비전 작업에서의 범용성은 향후 더 많은 데이터셋에서 검증되어야 한다.

실용적 활용

Meta-SpikeFormer는 저전력 비전 처리가 필요한 IoT, 모바일, 로봇 분야에서 활용 가능하다. 특히, event-based 센서와 결합하여 실시간 영상 처리 및 객체 인식에 사용할 수 있으며, neuromorphic 칩 설계에 기여하여 차세대 AI 하드웨어 개발을 촉진할 수 있다.