視覺生成中文本條件化的擴展性質
Scaling Properties of Text Conditioning in Visual Generation
July 31, 2026
作者: Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan
cs.AI
摘要
我們研究了視覺生成中文本條件化的實證規模特性。此類特性鮮少被度量,因為擴散損失不隨自然語言提示中的詞元數量而規模化。令人意外的是,我們發現收斂後的擴散損失會隨提示中結構化語言的數量而規模化。為了量化結構化語言,我們採用了兩種互補的度量:一種白箱似然度指標(GPG)與一種黑箱屬性指標(ED)。在一系列受控的訓練運行中,收斂後的擴散損失隨 GPG 近似線性下降,並隨 ED 呈冪律下降。在上述規模特性的引導下,我們透過建構帶有從影像中提取之語義與幾何標註的結構化提示來提升可擴散性,並透過監督式微調、冷啟動與驗證器門控的同策略蒸餾訓練提示器來提升可提示性。由此產生的系統在幾乎所有組合性、推理與世界知識基準上均勝過所有受評估的開放權重模型,同時在多數評估中達到或超越最強的封閉權重模型。
English
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.