ムービング・アルファベット:テキストからビデオ生成のための訓練データに関する対照研究
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
July 21, 2026
著者: Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin
cs.AI
要旨
テキストから動画を生成する技術は、過去5年間でモデルサイズ、データ、計算能力のスケーリングにより大幅に進歩した。モデルアーキテクチャとは異なり、学習データはしばしば十分に研究されていない。実世界のデータキュレーションは複雑かつ容易ではなく、生の動画からのクリップ選択や、テキストから動画へのマッピングを学習するための動画とテキストのペアを作成するキャプション付与が関わる。本研究では、データ分布とキャプションの品質がテキストから動画を生成するモデルにどのような影響を与えるかを調査する。制御された実験を可能にするため、我々はMoving Alphabetを導入する。これは、黒い背景に対して様々なフォント、色、サイズ、位置の文字を異なる方向と速度で動かすプロシージャルなテストベッドである。この設計により、真値のメタデータを劣化させることでデータ分布とキャプションの品質を精密に制御できる。実験から以下の三つの知見が得られた。(a) 動画の内容と時間長に関する多様でバランスの取れた分布が汎化のために重要であること。(b) キャプションの品質はモデルの性能と学習効率の両方に大きな影響を与え、テキストから動画を生成するモデルが動画理解能力に制約されていることを示唆する。(c) 分類器不要のガイダンスと高品質データでのファインチューニングにより、劣化したキャプションで学習されたモデルからの部分的な回復が可能であるが、質の低い事前学習データを完全に補償することはできない。これらの知見は大規模なテキストから動画を生成するモデルの開発に役立つと考えられ、事前学習データの科学により多くの注意を向けることを提唱する。
English
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.