ChatPaper.aiChatPaper

文本到图像像素空间扩散模型训练的实证研究

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

August 17, 2026
作者: Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi
cs.AI

摘要

本文研究生成建模中一个日益重要的话题:像素空间扩散模型。尽管已有大量研究探讨该话题,但大多数工作集中于小型或类别条件设置。因此,训练能够匹敌或超越成熟潜在空间对应模型的像素空间模型的实用方案仍然难以确定。通过全面的实证研究,我们首先观察到,直接在像素空间进行大规模预训练的收敛速度显著慢于潜在空间。这一观察促使我们设计一种从潜在空间到像素空间的策略:在潜在空间中高效获取生成先验,并在后训练阶段过渡到像素空间。随后,我们系统研究了决定这一过渡的关键设计选择,包括权重初始化、数据构成、预测目标、解码器架构和噪声调度,并确定了一种实用方案。该方案使所得像素空间模型在性能上匹敌或超越潜在空间对应模型,同时实现3.18至4.75倍的端到端推理加速。我们希望这些发现能为未来像素空间生成研究提供有用的实证洞见和实用指导。
English
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.