픽셀 공간 텍스트-이미지 확산 모델 훈련에 대한 실증 연구
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
August 17, 2026
저자: Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi
cs.AI
초록
본 논문은 생성 모델링에서 점점 더 중요해지고 있는 주제인 픽셀 공간 확산 모델을 연구한다. 다양한 연구들이 이 주제를 탐구해 왔지만, 대부분은 소규모 또는 클래스 조건부 설정에 초점을 맞추고 있다. 그 결과, 확립된 잠재 공간 모델들과 대등하거나 이를 능가하는 픽셀 공간 모델을 훈련시키기 위한 실용적인 방법론은 여전히 확립되지 않은 상태이다. 종합적인 실증 연구를 통해, 우리는 먼저 픽셀 공간에서의 직접적인 대규모 사전 학습이 잠재 공간에서보다 상당히 느리게 수렴한다는 점을 관찰한다. 이러한 관찰은 잠재 공간에서 생성적 사전 지식을 효율적으로 획득하고 사후 학습 단계에서 픽셀 공간으로 전환하는 잠재-픽셀 전략을 동기부여한다. 이후 우리는 가중치 초기화, 데이터 구성, 예측 대상, 디코더 아키텍처, 노이즈 스케줄 등 이 전환을 결정짓는 핵심 설계 선택들을 체계적으로 조사하고, 결과적으로 생성된 픽셀 공간 모델이 잠재 공간 모델과 대등하거나 이를 능가하면서 3.18배에서 4.75배의 종단 간 추론 속도 향상을 제공하는 실용적인 방법론을 식별한다. 우리의 발견이 픽셀 공간 생성에 관한 향후 연구에 유용한 실증적 통찰력과 실용적 지침을 제공할 수 있기를 기대한다.
English
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.