MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

S. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, S Ramaneswaran, Oriol Nieto, R. Duraiswami, Sreyan Ghosh, Dinesh Manocha

arXiv:2410.19168 · 2026-07-27 공개 · arXiv · PDF

large-language-models benchmark-evaluation reasoning-tasks multimodal-benchmark speech-recognition audio-language-models information-extraction audio-understanding

Abstract

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.

한국어 요약

한 줄 요약

MMAU는 10,000개의 음성-질의-답변 쌍을 포함한, 음성 이해 및 추론 능력을 평가하는 새로운 벤치마크로, 최신 모델들이 53% 미만의 정확도를 기록하며 기술적 한계를 드러낸다.

핵심 기여도

핵심 아이디어

MMAU는 단순한 음성 인식을 넘어, 전문가 수준의 지식과 복잡한 추론을 요구하는 음성 이해를 평가하기 위해 설계되었다. 기존 벤치마크는 주로 기본적인 인식 작업에 집중하지만, MMAU는 정보 추출과 추론 질문을 통해 모델이 27가지 구체적 기술을 보여야 한다. 예를 들어, 다중 화자의 역할 매핑, 감정 변화 탐지, 시간적 음향 이벤트 분석 등이 포함된다. 이는 모델이 단순히 음성을 인식하는 것을 넘어, 문맥과 지식을 바탕으로 추론해야 한다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

MMAU는 음성-언어 모델의 진정한 이해와 추론 능력을 평가하는 데 중요한 기준을 제공하며, 음성 이해 분야의 연구를 촉진할 것으로 기대된다. 그러나 MMAU는 인간 주석 데이터에 의존하므로, 주석 오류나 편향이 있을 수 있다. 또한, 일부 모델은 오디오 입력 처리에 어려움을 겪는 것으로 나타나, 모델 아키텍처와 학습 전략의 개선이 필요하다.

실용적 활용

MMAU는 음성 인식, 환경 소리 분석, 음악 이해 등 다양한 응용 분야에서 모델의 진정한 능력을 평가하는 데 활용될 수 있다. 특히, 대화형 AI, 보안 감시, 의료 음성 분석 등 전문가 수준의 추론이 필요한 분야에서 모델 개선을 촉진할 수 있다.