Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin

arXiv:2607.18789 · 2026-07-26 공개 · arXiv · PDF

video-understanding data-curation pre-training text-to-video training-data data-distribution large-scale-models classifier-free-guidance

Abstract

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.

한국어 요약

한 줄 요약

Moving Alphabet이라는 제어된 실험 환경을 통해 텍스트-비디오 생성 모델의 훈련 데이터 분포와 캡션 품질의 영향을 정량적으로 분석했다.

핵심 기여도

핵심 아이디어

Moving Alphabet은 텍스트-비디오 생성 모델의 훈련 데이터 품질을 연구하기 위한 제어된 환경으로, 다양한 폰트, 색상, 크기, 위치, 이동 방향과 속도를 가진 알파벳을 렌더링하여 생성된 비디오에 정확한 메타데이터를 부여한다. 이를 통해 데이터 분포와 캡션 품질을 인위적으로 조절할 수 있다. 이는 실제 데이터셋에서 관찰하기 어려운 요소들을 실험적으로 분리해 분석할 수 있는 기반을 제공한다. 핵심 통찰은 훈련 데이터의 질이 모델 성능에 직접적인 영향을 미친다는 점이다. 특히, 캡션의 정확도(정밀도)가 모델의 훈련 효율성과 최종 성능에 더 큰 영향을 미친다는 점이 강조된다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 텍스트-비디오 생성 모델의 학습 데이터 품질이 모델 성능에 직접적인 영향을 미친다는 점을 명확히 입증하며, 훈련 데이터 과학의 중요성을 강조한다. 특히, 캡션 정밀도가 모델 훈련 효율성과 최종 성능에 더 큰 영향을 미친다는 점은 기존 연구에서 다루어지지 않았던 핵심 발견이다. 그러나 Moving Alphabet은 제어된 환경에서의 실험으로, 실제 비디오 데이터와는 차이가 있을 수 있다. 또한, 이 연구는 비디오 콘텐츠의 복잡도와 클립 길이만 다루었으며, 해상도, 레이아웃 등 다른 요소는 포함되지 않았다.

실용적 활용

이 연구는 대규모 텍스트-비디오 생성 모델의 훈련 데이터 큐레이션 전략에 실질적인 가이드라인을 제공한다. 특히, 캡션 정밀도를 우선시하고, 다양한 복잡도와 길이의 비디오를 균형 있게 수집하는 것이 중요하다는 점이 강조된다. 이는 콘텐츠 생성, 광고, 엔터테인먼트 등 다양한 산업에서 텍스트 기반 비디오 생성 기술의 개선에 기여할 수 있다.