Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval.
We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts.
On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss.
Together, source mapping and input interventions reveal what drives retrieval.
한 줄 요약
MEG 신호에서 인지된 음성 재구성을 위해 해석 가능한 디코더를 설계하고, 그 기저 뇌 활성과 자극 특성을 분석한다.
핵심 기여도
- **공간 주의 모듈을 구면 조화함수 기반으로 재설계**하여 MEG 헬멧의 3차원 구조에 맞춤화.
- **270개의 브랜치를 25개로 축소**하고, 각 브랜치에 **시간 필터 추가**하여 뇌 신호의 시공간적 특성을 매칭.
- **MEG-MASC 데이터셋에서 39.75 ± 0.34%의 Top-1 정확도** 달성, 기존 모델 대비 약 20배 적은 파라미터 사용.
- **15개의 자극 특성**이 재구성에 기여하며, 특히 **침묵, 소리 강도, 모음, 음향 시작점**이 가장 영향력 있음.
핵심 아이디어
기존의 MEG-음성 디코딩 모델은 뇌 활성 신호를 wav2vec 2.0 임베딩으로 매핑하지만, 그 과정에서 뇌 해부학적 위치나 신경 활성의 동역학을 반영하지 못한다. 본 연구는 **물리적 및 생리학적 제약을 반영한 디코더 구조**를 제안한다. 구체적으로, **2D 푸리에 기반의 공간 주의 모듈을 구면 조화함수 기반으로 대체**하여 MEG 센서의 3차원 구조에 맞게 설계한다. 또한, **각 브랜치에 시간 필터를 추가**하여 뇌 신호의 시공간적 특성을 매칭하도록 한다. 이는 디코더가 학습한 가중치를 **소스 공간(source space)**에 매핑하여, 뇌 활성의 위치와 주파수 특성을 해석할 수 있게 한다. 특히, **왼쪽 뇌에 위치한 브랜치가 고주파 리듬 성분을 반영**한다는 점에서, 음성 인지 네트워크와의 일관성을 보인다.
기술적 접근법
- **입력 전처리**: 훈련 전에 **눈 깜빡임과 심장 신호 성분 제거**.
- **공간 주의 모듈**: 2D 푸리에 기반 → **구면 조화함수(spherical harmonics)** 기반으로 변경.
- **브랜치 수축**: 270개 → **K = 25개**로 축소.
- **시간 필터 추가**: 각 브랜치에 **학습 가능한 시간 필터** 추가하여 뇌 신호의 시공간적 매칭.
- **디코더 구조**: **shallower convolutional decoder** 사용.
- **학습 목표**: CLIP-style objective를 사용하여 MEG 임베딩과 wav2vec 2.0 임베딩 간 유사도 최대화.
주요 결과
- **MEG-MASC 데이터셋에서 39.75 ± 0.34%의 Top-1 정확도** 달성 (1005개 후보 중).
- **기존 디코더 대비 약 20배 적은 파라미터** 사용.
- **wav2vec 임베딩을 약 12차원으로 축소**해도 정확도 유지.
- **강한 시간 압축은 성능 저하** (약 10% 이상 감소).
- **15/19 자극 특성**이 재구성에 기여. **침묵, 소리 강도, 모음, 음향 시작점**이 가장 큰 영향.
- **무작위 단어 목록 실험**에서, **구조 없는 자극은 재구성 정보가 적음**.
의의 및 한계
본 연구는 **해석 가능한 디코더**를 통해 MEG 신호에서 인지된 음성을 재구성하는 데 성공했으며, **소스 공간 매핑과 입력 간섭을 통해 뇌 활성의 기저 요소를 분석**할 수 있는 새로운 방법론을 제시한다. 특히, **wav2vec 임베딩의 차원 축소 가능성**은 모델의 효율성 향상에 기여할 수 있다. 그러나, **시간 압축 시 성능 저하**는 디코더가 시간적 세부 정보에 의존한다는 점을 시사하며, 이는 모델의 일반화 능력에 한계를 줄 수 있다. 또한, **다양한 음성 자극에 대한 일반화 가능성**은 추가 실험을 통해 검증이 필요하다.
실용적 활용
본 연구는 **비침습적 MEG 기반 음성 디코딩** 기술의 발전에 기여하며, **뇌 인터페이스, 언어 신경프로세스, 수술 중 언어 맵핑** 등에 활용 가능하다. 특히, **해석 가능한 디코더는 뇌 활성의 생리학적 기저를 이해하는 데 유용**하며, **의학적 진단 및 치료 전략 개발**에도 기초 자료로 활용될 수 있다.