TempCloze: Can Video-LLMs Identify the Missing Middle?

arXiv:2609.01515 · 2026-09-12 공개 · arXiv · PDF

benchmark-evaluation temporal-reasoning egocentric-videos video-llms temporal-alignment tempcloze video-cloze long-take-videos

Abstract

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.

한국어 요약

한 줄 요약

TempCloze는 영상 기반 대형 언어 모델의 시각적 시간 추론 능력을 평가하기 위한 새로운 비디오 클로즈 벤치마크이다.

핵심 기여도

핵심 아이디어

기존의 Video-LLMs 평가 방식은 자연어 기반 질문과 선택지를 통해 이루어지며, 이는 언어적 단서를 통한 우회적 해결 가능성을 열어준다. 이를 해결하기 위해, TempCloze는 영상의 시작과 끝 클립이 주어졌을 때, 중간 클립을 4가지 후보 중에서 선택하도록 요구하는 클로즈 형식을 도입한다. 이는 모델이 시각적 증거를 기반으로 시간적 구조를 추론하도록 유도한다. TempCloze는 Semantic(어떤 사건이 발생해야 하는지), Alignment(언제 발생해야 하는지), Progression(어떻게 전개되어야 하는지)의 세 차원에서 디스트랙터를 구성하며, 공유된 장면과 객체를 통해 외형적 단서를 최소화한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용