ChatPaper.aiChatPaper

움직이는 알파벳: 텍스트-비디오 생성을 위한 훈련 데이터의 통제된 연구

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

July 21, 2026
저자: Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin
cs.AI

초록

지난 5년간 모델 크기, 데이터 및 연산량의 확장을 통해 텍스트-비디오 생성 기술이 크게 발전했다. 모델 아키텍처와 달리 학습 데이터는 종종 충분히 탐구되지 않았다. 실제 데이터 큐레이션은 복잡하고 간단하지 않으며, 원시 비디오에서 클립을 선택하고 캡션을 생성하여 텍스트-비디오 매핑 학습을 위한 비디오-텍스트 쌍을 만드는 과정을 포함한다. 본 연구는 데이터 분포와 캡션 품질이 텍스트-비디오 모델에 미치는 영향을 조사한다. 통제된 실험을 위해, 다양한 폰트, 색상, 크기 및 위치의 글자가 검은 배경 위에서 서로 다른 방향과 속도로 움직이는 Moving Alphabet이라는 절차적 테스트베드를 도입한다. 이 설계는 실제 메타데이터를 손상시켜 데이터 분포와 캡션 품질을 정밀하게 제어할 수 있게 한다. 실험 결과는 세 가지 발견점을 제시한다: a) 비디오 내용과 길이의 다양하고 균형 잡힌 분포는 일반화에 중요하다; b) 캡션 품질은 모델 성능과 학습 효율성 모두에 유의미한 영향을 미치며, 이는 텍스트-비디오 모델이 비디오 이해 능력에 의해 제한됨을 시사한다; c) 분류기-자유 유도 및 고품질 데이터에 대한 미세 조정은 손상된 캡션으로 학습된 모델에서 부분적인 회복을 제공하지만, 열악한 사전 학습 데이터를 완전히 보상할 수는 없다. 이러한 통찰이 대규모 텍스트-비디오 모델 개발에 도움이 될 수 있다고 믿으며, 사전 학습 데이터의 과학에 더 많은 관심을 기울일 것을 제안한다.
English
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.