訓練像素空間文字生成圖像擴散模型的實證研究
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
August 17, 2026
作者: Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi
cs.AI
摘要
本文探討生成建模中一個日益重要的主題:像素空間擴散模型。儘管已有大量研究探討此主題,但多數聚焦於小規模或類別條件設定。因此,要訓練出能與成熟的潛在空間模型匹敵甚至超越的像素空間模型,其實用配方仍然難以捉摸。透過一項全面的實證研究,我們首先觀察到,在像素空間直接進行大規模預訓練,其收斂速度明顯比在潛在空間中慢得多。此觀察促使我們提出一種「從潛在到像素」的策略,先在潛在空間中有效獲取生成先驗,再於後期訓練中轉換至像素空間。接著,我們系統性地探討了影響此轉換的關鍵設計選擇,包括權重初始化、資料組成、預測目標、解碼器架構與雜訊排程,並找出一個實用配方,使所產生的像素空間模型能與潛在空間模型相當或表現更佳,同時實現 3.18 至 4.75 倍的端到端推論加速。我們希望這些發現能為未來像素空間生成的研究提供有用的實證見解與實用指引。
English
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.