ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Lin Chen, Xilin Wei, Jinsong Li, Xiao-wen Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, Jiaqi Wang

arXiv:2406.04325 · 2026-07-27 공개 · arXiv · PDF

video-generation text-to-video video-captioning video-benchmarks sharegpt4video captioning-models temporal-description large-video-language-models

Abstract

We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos with various lengths and sources, developed through carefully designed data filtering and annotating strategy. 2) ShareCaptioner-Video, an efficient and capable captioning model for arbitrary videos, with 4.8M high-quality aesthetic videos annotated by it. 3) ShareGPT4Video-8B, a simple yet superb LVLM that reached SOTA performance on three advancing video benchmarks. To achieve this, taking aside the non-scalable costly human annotators, we find using GPT4V to caption video with a naive multi-frame or frame-concatenation input strategy leads to less detailed and sometimes temporal-confused results. We argue the challenge of designing a high-quality video captioning strategy lies in three aspects: 1) Inter-frame precise temporal change understanding. 2) Intra-frame detailed content description. 3) Frame-number scalability for arbitrary-length videos. To this end, we meticulously designed a differential video captioning strategy, which is stable, scalable, and efficient for generating captions for videos with arbitrary resolution, aspect ratios, and length. Based on it, we construct ShareGPT4Video, which contains 40K high-quality videos spanning a wide range of categories, and the resulting captions encompass rich world knowledge, object attributes, camera movements, and crucially, detailed and precise temporal descriptions of events. Based on ShareGPT4Video, we further develop ShareCaptioner-Video, a superior captioner capable of efficiently generating high-quality captions for arbitrary videos...

한국어 요약

한 줄 요약

ShareGPT4Video는 GPT4V와 자체 개발된 Differential Sliding-Window Captioning을 기반으로 40K 고질량 영상-캡션 데이터와 4.8M 캡션을 생성하여 영상 이해 및 생성 모델 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 영상 캡션 생성 방식은 다중 프레임 입력이나 프레임 연결 방식을 사용하지만, 이는 시간적 혼란과 세부 정보 손실을 유발한다. 이를 해결하기 위해, 연구팀은 **Differential Sliding-Window Captioning (DiffSW)** 전략을 제안한다. 이 전략은 첫 프레임에 대한 캡션을 생성한 후, **2개 프레임을 슬라이딩 윈도우로 입력**하여 **프레임 간의 변화**를 기술하는 방식이다. 이는 시간적 변화를 정확히 포착하고, 전체 영상의 캡션을 구성할 수 있도록 한다. 또한, 입력 프레임 수가 일정하게 유지되므로 GPT4V의 세부 정보 처리 능력을 활용할 수 있으며, **임의 길이의 영상에도 확장 가능**하다. 이는 기존 방식에서 발생하는 시간 혼동 문제와 세부 정보 손실 문제를 극복하는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용