Abra:擴展擴散圖像訓練
Abra: Scaling Diffusion Image Training
August 18, 2026
作者: Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan
cs.AI
摘要
計算最優擴展定律引導著前沿語言模型的訓練,但在視覺生成領域仍在很大程度上尚未被探索。我們使用受控的流匹配 Transformer 家族 Abra,對文字到圖像擴散模型進行了系統性的擴展定律研究。我們在橫跨三個數量級的計算量(10^{19} 至 10^{22} FLOPs)上進行訓練,達到了遠超以往工作的計算預算。我們證明,擴散模型的擴展行為與語言模型同樣可預測,但要實現最優訓練所需的數據卻要多得多:計算最優性出現於每個參數約 200 個圖像 token,是 LLM 的 Chinchilla 計算最優配方的十倍。我們進一步表明,與語言模型不同,擴散模型對過度訓練具有穩健性;從業者應寧可增加數據,而非選擇更大的模型。最後,我們證明這種可預測性不僅適用於訓練損失,也延伸至生成質量指標、最優 CFG 設定、表徵質量,甚至訓練曲線的形狀——這些曲線會收斂到一個通用形式。
English
Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute (10^{19} to 10^{22} FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.