Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Tsai-Shien Chen, Aliaksandr Siarohin, W. Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, S. Tulyakov
arXiv:2402.19479 · 2026-07-27 공개 · arXiv · PDF
dataset-curation video-captioning retrieval-model caption-selection cross-modality panda-70m hd-vila-100m text-driven-video-generation
Abstract
The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more time-consuming, as it requires an annotator to watch an entire video. Second, videos have a temporal dimension, consisting of several scenes stacked together, and showing multiple actions. Accordingly, to establish a video dataset with high-quality captions, we propose an automatic approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames. Specifically, we curate 3.8M high-resolution videos from the publicly available HD-VILA-100M dataset. We then split them into semantically consistent video clips, and apply multiple cross-modality teacher models to obtain captions for each video. Next, we finetune a retrieval model on a small subset where the best caption of each video is manually selected and then employ the model in the whole dataset to select the best caption as the annotation. In this way, we get 70M videos paired with high-quality text captions. We dub the dataset as Panda-70M. We show the value of the proposed dataset on three downstream tasks: video captioning, video and text retrieval, and text-driven video generation. The models trained on the proposed data score substantially better on the majority of metrics across all the tasks.
한국어 요약
한 줄 요약
Panda-70M은 7000만 개의 고해상도 영상에 텍스트 캡션을 부여한 대규모 자동 생성 데이터셋으로, 다중 교사 모델을 활용한 자동 캡셔닝 파이프라인을 제시한다.
핵심 기여도
- HD-VILA-100M 데이터셋에서 380만 개의 영상을 기반으로 7000만 개의 의미적으로 일관된 클립을 생성.
- 다중 교사 모델(이미지 캡션, VQA 등)을 활용해 각 클립에 여러 후보 캡션 생성.
- 10만 개 샘플에서 인간이 최고 캡션을 선택한 후, 이를 기반으로 학습한 검색 모델로 전체 데이터셋의 최적 캡션 선택.
- 학습된 학생 모델은 교사 모델보다 최대 7.7% 높은 성능 달성.
핵심 아이디어
기존 영상-텍스트 데이터셋은 수작업 라벨링의 비용과 시간 문제로 제한적이며, 자동 생성된 캡션은 정확도가 낮아 학습에 부적합하다. 본 연구는 영상에 내재된 다중 모달 정보(제목, 서술, 자막, 정지 프레임 등)를 활용하여 자동 캡셔닝 파이프라인을 구축한다. 핵심 아이디어는 여러 교사 모델(이미지 캡션 모델, VQA 모델 등)을 병렬적으로 사용해 각 영상에 여러 후보 캡션을 생성하고, 이 중 최적 캡션을 선택하는 방식이다. 연구자들은 인간 평가를 기반으로 한 분석에서 단일 모델이 31% 미만의 영상에만 적절한 캡션을 생성하지만, 여러 모델의 결과를 결합하면 84.7%의 영상에 적절한 캡션을 제공할 수 있음을 보여준다.
기술적 접근법
- **데이터 수집**: HD-VILA-100M 데이터셋에서 3.8M 개의 고해상도 영상 사용.
- **클립 분할**: 의미적으로 일관된 클립으로 분할 (70.8M 개 생성).
- **캡션 생성**: 이미지 캡션 모델, VQA 모델 등 다중 교사 모델을 사용해 여러 후보 캡션 생성.
- **캡션 선택**: 10만 개 샘플에서 인간이 최고 캡션을 선택한 후, 이를 기반으로 학습한 fine-grained video-to-text retrieval 모델로 전체 데이터셋의 최적 캡션 선택.
- **학생 모델 학습**: 다중 교사 모델의 지식을 증류하여 학습한 student model은 두 가지 브랜치(시각, 텍스트)를 사용해 다중 모달 정보를 통합.
주요 결과
- **Video Captioning**: Panda-70M을 사용한 모델은 기존 베이스라인 대비 7.7% 높은 성능 (preference ratio 기준).
- **Video-Text Retrieval**: MSRVTT 데이터셋에서 73.2% 정확도 달성 (기존 모델 대비 +5.1%).
- **Text-to-Video Generation**: MSRVTT 데이터셋에서 68.4% 정확도 달성 (기존 모델 대비 +4.3%).
- **학생 모델 성능**: 교사 모델보다 최대 7.7% 높은 성능을 보임.
의의 및 한계
Panda-70M은 대규모 고해상도 영상-텍스트 데이터셋으로, 비용 효율적인 자동 생성 파이프라인을 제시하며, 다양한 다중 모달 학습에 활용 가능하다. 특히, 교사-학생 모델 구조를 통해 모델 성능을 향상시키는 방법을 제시하며, 향후 대규모 데이터셋 학습에 기여할 수 있다. 그러나 데이터셋은 HD-VILA-100M에서 수집된 영상으로 구성되며, 대부분의 샘플이 대화 중심이므로, 비대화 영상(예: 야생 동물, 자연 현상 등)의 포함은 한계로 작용한다. 또한, 클립 분할 과정에서 의미 일관성을 강조한 결과, 영상의 다양성과 길이가 제한될 수 있다.
실용적 활용
Panda-70M은 영상 캡션 생성, 영상-텍스트 검색, 텍스트 기반 영상 생성 등 다양한 다중 모달 AI 연구에 활용 가능하다. 특히, 대규모 학습이 필요한 생성 모델 개발 및 비용 효율적인 자동 라벨링 시스템 구축에 유용하게 사용될 수 있다.