Abra: 拡散画像学習のスケーリング
Abra: Scaling Diffusion Image Training
August 18, 2026
著者: Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan
cs.AI
要旨
計算最適スケーリング則は最先端の言語モデルの訓練を導く一方、視覚生成においてはほぼ未解明のままである。我々は、制御されたフローマッチングトランスフォーマーのファミリーであるAbraを用いて、テキストから画像を生成する拡散モデルに対する系統的なスケーリング則研究を提示する。Abraは3桁にわたる計算量(10^19〜10^22 FLOPs)で訓練され、従来研究を大幅に上回る計算予算に達している。我々は、拡散モデルが言語モデルと同様に予測可能にスケールする一方、最適な訓練にははるかに多くのデータを必要とすることを実証する。すなわち、計算最適性はパラメータ当たり約200画像トークンで達成され、これはLLMに対するChinchillaの計算最適処方の10倍に相当する。さらに、言語モデルとは異なり、拡散モデルは過学習に対して頑健であり、実務者はより大きなモデルではなくより多くのデータを選ぶべきであることを示す。最後に、この予測可能性は訓練損失だけでなく、生成品質指標、最適なCFG設定、表現品質、さらには訓練曲線の形状にも及ぶことを示す。訓練曲線は普遍的な形に収束する。
English
Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute (10^{19} to 10^{22} FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.