VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng

arXiv:2608.14718 · 2026-08-18 공개 · arXiv · PDF

video-understanding multimodal-llms agentic-ai multi-turn tool-augmented mllm-evaluation real-world-scenarios external-tools

Abstract

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.

한국어 요약

한 줄 요약

VideoGAIA는 MLLM의 다단계, 도구 활용형 비디오 이해 능력을 평가하는 새로운 벤치마크로, 기존 90% 정확도에 가까운 단일 질문 대답 평가를 넘어서는 새로운 패러다임을 제시한다.

핵심 기여도

핵심 아이디어

VideoGAIA는 단일 질문 대답 중심의 비디오 이해 평가에서 벗어나, 비디오 인지, 외부 도구 사용, 다중 턴 간 정보 통합을 요구하는 새로운 평가 패러다임을 제시한다. 이는 비디오 에이전트가 단순히 질문에 답하는 것을 넘어, 비디오를 증거로 활용하여 실제 세계 문제를 해결할 수 있는 능력을 평가하는 것을 목표로 한다.

이를 위해, VideoGAIA는 인간 전문가와 모델이 협업하여 설계한 271개의 비디오-질문-답변 인스턴스를 포함하며, 각 인스턴스는 3명의 전문가에 의해 검증되어 정확성과 적절한 난이도를 보장한다. 평가 시에는 ReAct 프레임워크를 기반으로, 모델이 웹 검색, 페이지 방문, 비디오 재검토 등의 도구를 활용하여 정보를 수집하고 통합하는 과정을 평가한다.

기술적 접근법

주요 결과

의의 및 한계

VideoGAIA는 단일 질문 대답 중심의 비디오 이해 평가에서 벗어나, 비디오를 증거로 활용하는 에이전트 중심의 평가 체계를 제시함으로써, MLLM의 진정한 비디오 이해 능력을 평가할 수 있는 새로운 기준을 제시한다. 특히, 271개의 고질량 인스턴스와 ReAct 기반 평가 프로토콜은 비디오 에이전트 연구의 발전에 기여할 것으로 기대된다.

그러나, VideoGAIA는 아직 초기 단계의 벤치마크이며, 도구 사용의 효과성, 비디오-텍스트 간의 일관성, 모델의 장기적 추론 능력 등에 대한 추가 연구가 필요하다. 또한, 도구 사용 시 모델이 제공하는 요약 정보의 질이 최종 성능에 영향을 줄 수 있다는 점도 한계로 지적된다.

실용적 활용

VideoGAIA는 비디오 에이전트가 실제 세계 문제를 해결하는 능력을 평가하는 데 활용될 수 있으며, 특히 교육, 엔터테인먼트, 보안, 의료 등 비디오 기반 정보를 활용하는 산업 분야에서 모델의 신뢰성과 실용성을 평가하는 데 유용할 수 있다. 또한, MLLM의 도구 활용 능력과 장기적 추론 능력을 개선하기 위한 연구에도 기여할 수 있다.