Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Arushi Goel, Sreyan Ghosh, JaeHyeon Kim, Sonal Kumar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, R. Duraiswami, Dinesh Manocha, Rafael Valle, Bryan Catanzaro

arXiv:2507.08128 · 2026-08-15 공개 · arXiv · PDF

curriculum-training audio-understanding audio-language-model af-whisper long-audio multi-turn-chat audio-skills-xl voice-to-voice

Abstract

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.

한국어 요약

한 줄 요약

Audio Flamingo 3(AF3)는 음성, 소리, 음악을 통합한 대규모 오픈 소스 오디오-언어 모델로, 10분 긴 오디오 이해 및 추론 등 새로운 기능을 도입하고 20개 이상의 벤치마크에서 최고 성능을 달성했다.

핵심 기여도

핵심 아이디어

기존 대형 오디오-언어 모델(LALMs)은 주로 짧은 오디오 데이터로 학습되어 복잡한 추론 능력이 부족하다. AF3는 이 문제를 해결하기 위해 **AF-Whisper**라는 단일 오디오 인코더를 도입하여 음성, 소리, 음악을 통합적으로 학습한다. 이는 기존의 CLAP(소리/음악) + Whisper-v3(음성)와 비교해 더 높은 효율성과 성능을 제공한다. 또한, **AudioSkills-XL**과 같은 대규모 데이터셋을 통해 모델이 다양한 추론 상황에 노출되도록 유도하며, **5단계 커리큘럼 학습**을 통해 점진적으로 복잡도를 증가시켜 모델의 추론 능력을 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

AF3는 오디오-언어 모델 분야에서 **전면 오픈 소스**라는 점에서 혁신적이며, 모델 가중치, 데이터셋, 코드를 모두 공개함으로써 연구 투명성과 재현성을 높였다. 특히, **AudioSkills-XL**, **AF-Think** 등 새로운 데이터셋과 **AF-Whisper** 인코더는 오디오-언어 모델의 성능 향상에 중요한 기여를 했다. 그러나, **Gemini**와 같은 클로즈드 소스 모델과 비교하면 일부 벤치마크에서 여전히 성능 격차가 존재하며, **음성-음성 상호작용 평가**는 범위를 넘어서므로 포함되지 않았다.

실용적 활용

AF3는 음성 비서, 멀티모달 챗봇, 오디오 분석 시스템 등에 적용 가능하다. 특히, **다중 턴 대화**와 **긴 오디오 추론** 기능은 고객 지원, 교육, 엔터테인먼트 분야에서 활용도가 높으며, **오픈 소스 공개**를 통해 연구자와 개발자들이 모델을 자유롭게 활용하고 개선할 수 있다.