Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S. Sakshi, JaeHyeon Kim, Wei Ping, Rafael Valle, Dinesh Manocha, Bryan Catanzaro, AF-CLAP Contrastive Loss

arXiv:2503.03983 · 2026-07-27 공개 · arXiv · PDF

curriculum-learning audio-language-model small-language-model reasoning-abilities long-audio longaudio-dataset clap-model longaudiobench

Abstract

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Language Model (ALM) with advanced audio understanding and reasoning capabilities. AF2 leverages (i) a custom CLAP model, (ii) synthetic Audio QA data for fine-grained audio reasoning, and (iii) a multi-stage curriculum learning strategy. AF2 achieves state-of-the-art performance with only a 3B parameter small language model, surpassing large open-source and proprietary models across over 20 benchmarks. Next, for the first time, we extend audio understanding to long audio segments (30 secs to 5 mins) and propose LongAudio, a large and novel dataset for training ALMs on long audio captioning and question-answering tasks. Fine-tuning AF2 on LongAudio leads to exceptional performance on our proposed LongAudioBench, an expert annotated benchmark for evaluating ALMs on long audio understanding capabilities. We conduct extensive ablation studies to confirm the efficacy of our approach. Project Website: https://research.nvidia.com/labs/adlr/AF2/.

한국어 요약

한 줄 요약

Audio Flamingo 2는 3B 파라미터 소형 언어 모델로, 20개 이상의 벤치마크에서 최고 성능을 달성한 다기능 오디오-언어 모델이다.

핵심 기여도

핵심 아이디어

Audio Flamingo 2는 기존 오디오-언어 모델이 전문가 수준의 추론 능력 부족과 긴 오디오 처리 능력 부재라는 두 가지 주요 한계를 해결하고자 설계되었다. 이를 위해, AF2는 CLAP 기반의 AF-CLAP 인코더를 통해 더 높은 품질의 오디오 표현을 구축하고, 합성 데이터를 포함한 8M 개 이상의 오디오-캡션 쌍으로 학습한다. 또한, AudioSkills라는 합성 QA 데이터셋을 통해 세부적인 추론 능력을 향상시키며, 3단계 과정 학습 전략을 통해 모델 성능을 점진적으로 향상시킨다. LongAudio 데이터셋을 통해 5분 길이의 오디오를 처리할 수 있는 능력을 확장함으로써, 기존 연구가 다루지 못한 새로운 범주를 개척한다.

기술적 접근법

주요 결과

의의 및 한계

Audio Flamingo 2는 소형 모델로도 전문가 수준의 오디오 추론과 긴 오디오 처리를 가능하게 함으로써, 모델 크기와 성능 간의 관계를 재정의하는 중요한 기여를 한다. LongAudio와 LongAudioBench는 오디오-언어 연구 분야에 새로운 평가 기준과 학습 자원을 제공한다. 그러나, 합성 데이터에 의존하는 점과 전문가 주석 데이터의 부족은 모델의 일반화 능력에 영향을 줄 수 있는 한계로 작용할 수 있다.

실용적 활용

Audio Flamingo 2는 산업 환경에서의 이상 탐지, 감정 인식, 장애인 지원 기술 등 다양한 분야에 적용 가능하다. 특히, 긴 오디오를 처리할 수 있는 능력은 보안, 교육, 헬스케어 분야에서 실용적 활용이 기대된다.