TempCompass: Do Video LLMs Really Understand Videos?

Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, Lu Hou

arXiv:2403.00476 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation video-llms video-analysis temporal-perception llm-based-evaluation instruction-generation temporal-aspects temporal-bias

Abstract

Recently, there is a surge in interest surrounding video large language models (Video LLMs). However, existing benchmarks fail to provide a comprehensive feedback on the temporal perception ability of Video LLMs. On the one hand, most of them are unable to distinguish between different temporal aspects (e.g., speed, direction) and thus cannot reflect the nuanced performance on these specific aspects. On the other hand, they are limited in the diversity of task formats (e.g., only multi-choice QA), which hinders the understanding of how temporal perception performance may vary across different types of tasks. Motivated by these two problems, we propose the \textbf{TempCompass} benchmark, which introduces a diversity of temporal aspects and task formats. To collect high-quality test data, we devise two novel strategies: (1) In video collection, we construct conflicting videos that share the same static content but differ in a specific temporal aspect, which prevents Video LLMs from leveraging single-frame bias or language priors. (2) To collect the task instructions, we propose a paradigm where humans first annotate meta-information for a video and then an LLM generates the instruction. We also design an LLM-based approach to automatically and accurately evaluate the responses from Video LLMs. Based on TempCompass, we comprehensively evaluate 8 state-of-the-art (SOTA) Video LLMs and 3 Image LLMs, and reveal the discerning fact that these models exhibit notably poor temporal perception ability. Our data will be available at https://github.com/llyx97/TempCompass.

한국어 요약

한 줄 요약

TempCompass는 Video LLM의 시간 인지 능력을 체계적으로 평가하기 위한 새로운 벤치마크로, 8개의 최신 Video LLM이 시간 인지 능력에서 부족함을 드러냄.

핵심 기여도

핵심 아이디어

기존 벤치마크는 Video LLM의 시간 인지 능력을 정확히 평가하지 못했다. 특히, 시간의 다양한 측면(예: 속도, 방향)을 구분하지 못하고, 태스크 형식이 단일(예: 단일 선택 QA)이라 평가의 완전성을 해쳤다. TempCompass는 이 문제를 해결하기 위해 시간 측면과 태스크 형식의 다변화를 도입했다.

핵심 아이디어는 ‘상호 배타적 비디오’를 구성하는 것이다. 같은 정적 콘텐츠를 공유하지만 특정 시간 측면(예: 속도, 방향)에서만 차이를 두어, Video LLM이 단일 프레임 정보나 언어 편향에 의존하지 않고 시간적 변화를 인식해야 한다는 점을 강제한다. 이는 Video LLM이 진정한 시간 인지 능력을 갖추고 있는지 평가하는 데 효과적이다.

기술적 접근법

주요 결과

의의 및 한계

TempCompass는 Video LLM의 시간 인지 능력을 체계적으로 평가할 수 있는 첫 번째 벤치마크로, Video LLM이 단일 프레임 편향이나 언어 편향에 의존하는 문제를 드러냈다. 또한, 다양한 태스크 형식을 도입함으로써 평가의 완전성을 높였다.

하지만, TempCompass는 Shutterstock 플랫폼에서만 데이터를 수집했기 때문에, 다른 도메인에서의 일반화 가능성은 명시되지 않았다. 또한, 평가에 사용된 ChatGPT(gpt3.5-turbo)는 최신 모델이 아닌 점에서 한계가 있을 수 있다.

실용적 활용

TempCompass는 Video LLM의 시간 인지 능력을 개선하기 위한 연구 개발에 활용될 수 있으며, 특히 영상 기반 교육, 자율 주행, 보안 감시 등 시간적 변화를 이해해야 하는 산업 분야에서 모델 평가에 유용할 수 있다.