Abra:扩展扩散图像训练
Abra: Scaling Diffusion Image Training
August 18, 2026
作者: Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan
cs.AI
摘要
计算最优缩放定律指导着前沿语言模型的训练,但在视觉生成领域仍鲜有探索。我们利用Abra开展了一项系统的缩放定律研究——这是一个受控的流匹配变换器系列,训练计算量跨越三个数量级(10^19至10^22 FLOPs),远超以往工作所达到的计算预算。我们证明,扩散模型的缩放规律与语言模型同样可预测,但实现最优训练所需的数据量要大得多:计算最优性出现在每参数约200个图像token处,是LLM的Chinchilla计算最优配方的十倍。我们还表明,与语言模型不同,扩散模型对过度训练具有鲁棒性,实践者应优先考虑更多数据而非更大模型。最后,我们展示了这种可预测性不仅体现在训练损失上,还延伸至生成质量指标、最优CFG设置、表示质量,甚至训练曲线的形状——这些曲线会坍缩为一种通用形式。
English
Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute (10^{19} to 10^{22} FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.