AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, Jingren Zhou

arXiv:2402.07729 · 2026-07-27 공개 · arXiv · PDF

large-language-models instruction-following evaluation-framework gpt-4 audio-language-models asr generative-comprehension audio-signal-processing

Abstract

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded advancements in this field. Previous models primarily focus on assessing different fundamental tasks, such as Automatic Speech Recognition (ASR), and lack an assessment of the open-ended generative capabilities centered around audio. Thus, it is challenging to track the progression in the Large Audio-Language Models (LALMs) domain and to provide guidance for future improvement. In this paper, we introduce AIR-Bench (\textbf{A}udio \textbf{I}nst\textbf{R}uction \textbf{Bench}mark), the first benchmark designed to evaluate the ability of LALMs to understand various types of audio signals (including human speech, natural sounds, and music), and furthermore, to interact with humans in the textual format. AIR-Bench encompasses two dimensions: \textit{foundation} and \textit{chat} benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions, intending to inspect the basic single-task ability of LALMs. The latter one contains 2k instances of open-ended question-and-answer data, directly assessing the comprehension of the model on complex audio and its capacity to follow instructions. Both benchmarks require the model to generate hypotheses directly. We design a unified framework that leverages advanced language models, such as GPT-4, to evaluate the scores of generated hypotheses given the meta-information of the audio. Experimental results demonstrate a high level of consistency between GPT-4-based evaluation and human evaluation. By revealing the limitations of existing LALMs through evaluation results, AIR-Bench can provide insights into the direction of future research.

한국어 요약

한 줄 요약

AIR-Bench는 대규모 오디오-언어 모델의 생성적 이해 능력을 평가하는 첫 번째 벤치마크로, 19개의 기초 태스크와 2,000개의 개방형 질문을 포함한다.

핵심 기여도

핵심 아이디어

기존 연구는 오디오-언어 모델의 평가가 ASR 등 단일 태스크에 집중되어 있으며, 생성적 대화 능력을 평가하는 체계적인 벤치마크가 부재한 문제를 지적한다. AIR-Bench는 이 문제를 해결하기 위해, 다양한 오디오 신호(인간 음성, 자연 소리, 음악)를 포함한 *foundation* 및 *chat* 벤치마크를 설계하고, 모델이 직접 가설을 생성하도록 요구하는 평가 프레임워크를 제안한다. 특히, GPT-4를 활용한 평가 방식은 기존 자동 평가 지표(WER, ROUGE 등)보다 인간 판단과 더 높은 일관성을 보인다. 또한, 오디오 믹싱 전략을 통해 실제 상황에 가까운 복잡한 오디오를 생성하여 평가의 현실성을 높였다.

기술적 접근법

주요 결과

의의 및 한계

AIR-Bench는 대규모 오디오-언어 모델의 생성적 이해 능력을 체계적으로 평가할 수 있는 첫 번째 벤치마크로, 기존 ASR 중심 평가의 한계를 극복하고, 모델의 진정한 대화 능력을 평가할 수 있는 기반을 제공한다. 또한, GPT-4를 활용한 평가 프레임워크는 자동 평가의 객관성과 일관성을 높이는 데 기여한다. 그러나, GPT-4가 오디오 입력을 직접 받지 못하기 때문에, 오디오의 메타정보에 의존하는 한계가 존재하며, 이는 평가의 완전성을 제한할 수 있다. 또한, 현재는 9개의 모델만 평가되었으며, 더 많은 모델을 포함한 평가가 필요하다.

실용적 활용

AIR-Bench는 대규모 오디오-언어 모델의 개발 및 비교를 위한 표준 평가 도구로 활용될 수 있으며, 특히 음성 인식, 음악 이해, 자연 소리 분석 등 다양한 응용 분야에서 모델 성능을 객관적으로 평가하는 데 유용하다. 또한, 채팅형 응용(예: 음성 기반 가상 보조) 개발에도 중요한 기준이 될 수 있다.