SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang

arXiv:2608.17426 · 2026-08-20 공개 · arXiv · PDF

video-generation instruction-following vision-language-model evaluation-protocol outcome-achievement generation-reliability semantic-task-completion semcomp-bench

Abstract

We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.

한국어 요약

한 줄 요약

SemComp-Bench는 생성된 동영상이 지시된 결과를 달성하면서 참조 이미지와 의미적 연관성을 유지하는 능력을 평가하는 새로운 벤치마크 프로토콜이다.

핵심 기여도

핵심 아이디어

기존 비디오 생성 모델은 시각적 품질과 시간적 일관성을 중시하지만, 지시된 작업 결과를 달성하면서 참조 이미지와 의미적 연관성을 유지하는 능력은 부족하다. 이를 해결하기 위해, **Semantic Task Completion Video Generation**이라는 새로운 작업 개념을 제시한다. 이 작업은 생성된 동영상이 **고수준 의미적 연관**(semantic grounding)을 유지하면서 지시된 결과를 달성하는 것을 목표로 한다. 예를 들어, 지폐 이미지와 "지폐를 거북이 모양으로 접으세요"라는 지시가 주어졌을 때, 생성된 동영상은 지폐가 거북이 모양의 오리게임으로 변해야 하며, 중간 과정은 필요하지 않다. 이는 기존의 완전한 시퀀스 생성을 요구하는 작업과 구별된다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용