OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, Ying Tai

arXiv:2407.02371 · 2026-07-27 공개 · arXiv · PDF

video-generation diffusion-transformer multi-modal text-to-video ablation-study dataset-curation captioning openvid-1m

Abstract

Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previous popular video datasets, e.g. WebVid-10M and Panda-70M, are either with low quality or too large for most research institutions. Therefore, it is challenging but crucial to collect a precise high-quality text-video pairs for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of thoroughly extracting semantic information from text prompt. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT.

한국어 요약

한 줄 요약

OpenVid-1M은 100만 개 이상의 고품질 텍스트-비디오 쌍을 포함한 대규모 데이터셋으로, MVDiT라는 새로운 다중 모달 디퓨전 트랜스포머를 제안하여 텍스트-비디오 생성을 개선한다.

핵심 기여도

핵심 아이디어

기존 텍스트-비디오 생성 연구는 WebVid-10M, Panda-70M와 같은 대규모 데이터셋을 사용했지만, 이들은 품질이 낮거나 설명이 부족한 문제가 있었다. OpenVid-1M은 512×512 이상의 해상도를 가지며, LLaVA-v1.6-34b를 사용해 생성된 상세한 캡션을 포함하여 텍스트-비디오 생성의 품질을 높이는 데 기여한다. 또한, MVDiT는 기존 DiT 기반 모델과 달리 병렬적인 시각-텍스트 구조를 도입하여, 시각 토큰의 구조 정보와 텍스트 토큰의 의미 정보를 동시에 추출한다. 이는 기존의 단순 cross attention 모듈이 의미 정보를 충분히 활용하지 못하는 문제를 해결한다.

기술적 접근법

주요 결과

의의 및 한계

OpenVid-1M은 텍스트-비디오 생성 분야에서 대규모 고품질 데이터의 부족 문제를 해결하며, MVDiT는 텍스트 의미 정보를 더 잘 활용할 수 있는 새로운 아키텍처를 제시한다. 이는 향후 더 정교한 비디오 생성 모델 개발에 기여할 수 있다. 그러나 모델은 여전히 자연 장면의 복잡한 역학과 움직임을 정확히 모델링하는 데 어려움이 있으며, 이는 더 많은 고품질 데이터와 모델 확장을 통해 개선될 수 있다.

실용적 활용

OpenVid-1M은 연구 기관 및 기업에서 텍스트-비디오 생성 모델 개발에 활용할 수 있으며, 특히 고해상도 비디오 생성, 콘텐츠 자동화, VR/AR 분야에서 활용 가능하다. MVDiT는 텍스트 기반의 시각 콘텐츠 생성을 필요로 하는 산업에서 즉시 적용할 수 있는 기술적 기반을 제공한다.