GigaChat Audio: Time-aware Large Audio Language Model

Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

arXiv:2607.10387 · 2026-07-21 공개 · arXiv · PDF

long-context benchmark-evaluation temporal-grounding synthetic-supervision audio-llm model-release time-aware duration-mixture

Abstract

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

한국어 요약

한 줄 요약

GigaChat Audio는 최대 120분의 오디오 입력을 처리하며, 시간 기반 질문에 정확한 타임스탬프를 포함한 답변을 생성하는 시간 인식 대형 오디오 언어 모델이다.

핵심 기여도

핵심 아이디어

기존 오디오 조건 대형 언어 모델(LLM)은 긴 오디오에서 시간 정합성(temporal grounding)이 불안정한 문제가 있었다. 이에 GigaChat Audio는 오디오 토큰 스트림에 주기적으로 시간 마커(inter-timings)를 삽입하여 시간 정보를 명시적으로 표현하는 방식을 도입했다. 이는 시간 기반 질문에 정확한 타임스탬프를 생성할 수 있도록 하며, 특히 60초 간격으로 삽입하는 것이 실험적으로 효과적임을 보여준다. 또한, 학습 데이터는 초~시간 길이의 오디오를 포함한 다중 길이 혼합(mixture of durations)으로 구성되어 있어, 모델이 다양한 길이의 오디오에 일반화할 수 있도록 한다. 이는 단일 길이 학습에서 성능 저하가 발생하는 문제를 해결한다.

기술적 접근법

주요 결과

의의 및 한계

GigaChat Audio는 오디오-조건 LLM에서 시간 정합성 문제를 해결하는 첫 번째 시도로, 120분 길이의 오디오 처리 및 시간 기반 답변 생성 기능을 제공한다. 특히, 주기적 시간 마커와 다중 길이 학습 데이터의 중요성을 실증적으로 입증하여, 시간 인식 오디오 모델 연구에 기초를 제공한다. 그러나, 모델은 학습 데이터의 질과 양에 크게 의존하며, 실제 응용 시 시간 마커의 빈도 조절이 성능-계산 비용 간 트레이드오프를 요구한다. 또한, 시간 기반 요약 평가에서 자동 평가 지표와 인간 평가 간의 불일치가 일부 존재한다.

실용적 활용

GigaChat Audio는 회의, 팟캐스트, 강의, 고객센터 로그 등 긴 오디오 콘텐츠를 처리하는 인터랙티브 시스템에 적용 가능하다. 사용자는 질문을 통해 특정 시간대의 오디오를 탐색하거나, 시간 기반 요약을 생성할 수 있어, 정보 검색 및 분석 효율성을 높일 수 있다. 특히, 고객 지원, 교육, 미디어 산업에서 실용적 활용이 기대된다.