視覚生成におけるテキスト条件付けのスケーリング特性
Scaling Properties of Text Conditioning in Visual Generation
July 31, 2026
著者: Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan
cs.AI
要旨
我々は、視覚生成におけるテキスト条件付けの経験的スケーリング特性を研究する。このような特性は、拡散損失が自然言語プロンプト内のトークン数に応じてスケールしないため、これまでほとんど測定されてこなかった。驚くべきことに、収束した拡散損失はプロンプト内の構造化言語の量とともにスケールすることがわかった。構造化言語を定量化するために、我々はホワイトボックス尤度指標(GPG)とブラックボックス属性指標(ED)という相補的な2つの尺度を採用する。制御された訓練実行を通じて、収束した拡散損失はGPGに対してほぼ線形に減少し、EDに対してはべき乗則に従う。これらのスケーリング特性に導かれ、我々は画像から導出された意味的・幾何学的アノテーションを用いて構造化プロンプトを構築することで拡散容易性を改善し、教師ありファインチューニング、コールドスタート、検証器ゲート付きオン方策蒸留を通じてプロンプターを訓練することでプロンプト容易性を改善する。得られたシステムは、評価されたすべてのオープンウェイトモデルを、ほとんどすべての構成的・推論・世界知識ベンチマークで上回り、ほとんどの評価で最強のクローズドウェイトモデルに並ぶか、それを上回る。
English
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.