ピクセル空間テキストから画像への拡散モデルの学習に関する実証研究
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
August 17, 2026
著者: Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi
cs.AI
要旨
本論文では、生成モデリングにおいてますます重要性を増しているトピックである、ピクセル空間拡散モデルを調査する。このトピックを探求した研究は数多く存在するが、そのほとんどは小規模またはクラス条件付きの設定に焦点を当てている。その結果、確立された潜在空間モデルに匹敵するか、それを上回るピクセル空間モデルを訓練するための実用的な手法は、依然として確立されていない。包括的な実証研究を通じて、我々はまず、ピクセル空間での直接的な大規模事前学習は、潜在空間での学習よりもかなり遅く収束することを観察する。この観察は、潜在空間で生成事前知識を効率的に獲得し、ポストトレーニング中にピクセル空間へ移行する、潜在空間からピクセル空間への戦略を動機付ける。次に、我々はこの移行を左右する主要な設計選択肢(重みの初期化、データ構成、予測ターゲット、デコーダアーキテクチャ、ノイズスケジュールなど)を体系的に調査し、結果として得られるピクセル空間モデルが潜在空間モデルに匹敵するか、それを上回りつつ、エンドツーエンドの推論速度を3.18倍から4.75倍向上させる実用的な手法を特定する。我々の発見が、ピクセル空間生成に関する将来の研究にとって有益な実証的知見と実用的なガイドラインを提供することを期待している。
English
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.