ChatPaper.aiChatPaper

시각 생성에서 텍스트 조건화의 스케일링 속성

Scaling Properties of Text Conditioning in Visual Generation

July 31, 2026
저자: Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan
cs.AI

초록

우리는 시각 생성에서 텍스트 조건화의 경험적 스케일링 속성을 연구한다. 확산 손실은 자연어 프롬프트의 토큰 수에 따라 확장되지 않기 때문에 이러한 속성은 거의 측정된 바 없다. 놀랍게도, 수렴된 확산 손실은 프롬프트 내 구조화된 언어의 양에 따라 확장된다는 것을 발견한다. 구조화된 언어를 정량화하기 위해, 우리는 화이트박스 우도 지표(GPG)와 블랙박스 속성 지표(ED)라는 두 가지 상보적 측정 방식을 적용한다. 통제된 훈련 실행들에 걸쳐, 수렴된 확산 손실은 GPG에 대해 대략 선형적으로 감소하며 ED에 대해 멱법칙을 따른다. 이러한 스케일링 속성에 기반하여, 우리는 이미지에서 추출한 의미론적 및 기하학적 주석을 포함하는 구조화된 프롬프트를 구성함으로써 확산성을 개선하고, 지도 미세 조정, 콜드 스타트 및 검증기 기반 온정책 증류를 통해 프롬프터를 훈련함으로써 프롬프트성을 개선한다. 결과 시스템은 거의 모든 구성적, 추론 및 세계 지식 벤치마크에서 평가된 모든 오픈 가중치 모델을 능가하며, 대부분의 평가에서 가장 강력한 폐쇄 가중치 모델과 동등하거나 이를 능가한다.
English
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.