Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Tsai-Shien Chen, Aliaksandr Siarohin, W. Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, S. Tulyakov

arXiv:2402.19479 · 2026-07-27 공개 · arXiv · PDF

dataset-curation video-captioning retrieval-model caption-selection cross-modality panda-70m hd-vila-100m text-driven-video-generation

Abstract

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more time-consuming, as it requires an annotator to watch an entire video. Second, videos have a temporal dimension, consisting of several scenes stacked together, and showing multiple actions. Accordingly, to establish a video dataset with high-quality captions, we propose an automatic approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames. Specifically, we curate 3.8M high-resolution videos from the publicly available HD-VILA-100M dataset. We then split them into semantically consistent video clips, and apply multiple cross-modality teacher models to obtain captions for each video. Next, we finetune a retrieval model on a small subset where the best caption of each video is manually selected and then employ the model in the whole dataset to select the best caption as the annotation. In this way, we get 70M videos paired with high-quality text captions. We dub the dataset as Panda-70M. We show the value of the proposed dataset on three downstream tasks: video captioning, video and text retrieval, and text-driven video generation. The models trained on the proposed data score substantially better on the majority of metrics across all the tasks.

한국어 요약

한 줄 요약

Panda-70M은 7000만 개의 고해상도 영상에 텍스트 캡션을 부여한 대규모 자동 생성 데이터셋으로, 다중 교사 모델을 활용한 자동 캡셔닝 파이프라인을 제시한다.

핵심 기여도

핵심 아이디어

기존 영상-텍스트 데이터셋은 수작업 라벨링의 비용과 시간 문제로 제한적이며, 자동 생성된 캡션은 정확도가 낮아 학습에 부적합하다. 본 연구는 영상에 내재된 다중 모달 정보(제목, 서술, 자막, 정지 프레임 등)를 활용하여 자동 캡셔닝 파이프라인을 구축한다. 핵심 아이디어는 여러 교사 모델(이미지 캡션 모델, VQA 모델 등)을 병렬적으로 사용해 각 영상에 여러 후보 캡션을 생성하고, 이 중 최적 캡션을 선택하는 방식이다. 연구자들은 인간 평가를 기반으로 한 분석에서 단일 모델이 31% 미만의 영상에만 적절한 캡션을 생성하지만, 여러 모델의 결과를 결합하면 84.7%의 영상에 적절한 캡션을 제공할 수 있음을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

Panda-70M은 대규모 고해상도 영상-텍스트 데이터셋으로, 비용 효율적인 자동 생성 파이프라인을 제시하며, 다양한 다중 모달 학습에 활용 가능하다. 특히, 교사-학생 모델 구조를 통해 모델 성능을 향상시키는 방법을 제시하며, 향후 대규모 데이터셋 학습에 기여할 수 있다. 그러나 데이터셋은 HD-VILA-100M에서 수집된 영상으로 구성되며, 대부분의 샘플이 대화 중심이므로, 비대화 영상(예: 야생 동물, 자연 현상 등)의 포함은 한계로 작용한다. 또한, 클립 분할 과정에서 의미 일관성을 강조한 결과, 영상의 다양성과 길이가 제한될 수 있다.

실용적 활용

Panda-70M은 영상 캡션 생성, 영상-텍스트 검색, 텍스트 기반 영상 생성 등 다양한 다중 모달 AI 연구에 활용 가능하다. 특히, 대규모 학습이 필요한 생성 모델 개발 및 비용 효율적인 자동 라벨링 시스템 구축에 유용하게 사용될 수 있다.