ChatPaper.aiChatPaper

视觉生成中文本条件的缩放性质

Scaling Properties of Text Conditioning in Visual Generation

July 31, 2026
作者: Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan
cs.AI

摘要

我们研究了视觉生成中文本条件的经验缩放性质。此类性质很少被测量,因为扩散损失并不随自然语言提示中的词元数量而缩放。令人惊讶的是,我们发现收敛的扩散损失与提示中结构化语言的数量呈缩放关系。为了量化结构化语言,我们采用了两种互补的度量:白盒似然度量(GPG)和黑盒属性度量(ED)。在受控的训练运行中,收敛的扩散损失随GPG近似线性下降,并随ED呈幂律关系。在这些缩放性质的指导下,我们通过构建带有从图像中提取的语义和几何注释的结构化提示来提高可扩散性,并通过监督微调、冷启动和验证器门控的同策略蒸馏来训练提示器,从而提高可提示性。所得到的系统在几乎所有的组合、推理和世界知识基准上均优于所有评估过的开放权重模型,同时在大多数评估中达到或超越最强的封闭权重模型。
English
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.