ChatPaper.aiChatPaper

运动字母表:面向文本生成视频的训练数据受控研究

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

July 21, 2026
作者: Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin
cs.AI

摘要

文本到视频生成在过去五年中通过模型规模、数据和计算量的扩展取得了显著进展。与模型架构不同,训练数据往往未得到充分探索。现实世界的数据整理复杂且非琐碎,涉及从原始视频中挑选片段并进行字幕标注,以生成用于学习文本到视频映射的视频-文本对。我们研究了数据分布和字幕质量对文本到视频模型的影响。为进行受控实验,我们引入了动态字母表(Moving Alphabet),这是一个程序化测试平台,能够在黑色背景上以不同字体、颜色、大小和位置渲染字母,并以不同方向和速度移动。这一设计通过破坏真实元数据,实现了对数据分布和字幕质量的精确控制。我们的实验得出三项发现:a)视频内容和时长的多样性与均衡分布对泛化能力至关重要;b)字幕质量显著影响模型性能与训练效率,表明文本到视频模型的能力受限于视频理解水平;c)无分类器引导以及在高质量数据上的微调能够部分弥补在劣质字幕上训练的模型,但无法完全补偿预训练数据质量不足的问题。我们相信这些见解可为大规模文本到视频模型的开发提供参考,并呼吁对预训练数据的科学给予更多关注。
English
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.