移動字母:文字生成影片訓練數據的對照研究
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
July 21, 2026
作者: Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin
cs.AI
摘要
過去五年間,透過擴大模型規模、數據量與運算能力,文字生成影片技術已取得顯著進展。不同於模型架構,訓練數據的潛力往往未被充分探索。真實世界的數據策展過程複雜且不簡單,涉及從原始影片中選取片段並添加字幕,以建立用於學習文字與影片對應關係的影片-文字配對。我們研究數據分佈與字幕品質如何影響文字生成影片模型。為了進行受控實驗,我們引入「移動字母表」——一個程序化測試平台,能在黑色背景上生成不同字體、顏色、大小與位置的字母,並以不同方向與速度移動。此設計能透過破壞真實標註元數據,精準控制數據分佈與字幕品質。我們的實驗得出三項發現:a) 影片內容與時長的多樣化且均衡分佈,對泛化能力至關重要;b) 字幕品質顯著影響模型表現與訓練效率,暗示文字生成影片模型受限於影片理解能力;c) 無分類器引導與高品質數據微調,雖能部分補償基於劣質字幕訓練的模型,但無法完全彌補不佳的預訓練數據。我們認為這些見解有助於大規模文字生成影片模型的開發,並呼籲更重視預訓練數據的科學研究。
English
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.