GenFirst:安定的なエンドツーエンド潜在生成モデリングのための再構成に先立つ生成
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
August 29, 2026
著者: Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang
cs.AI
要旨
潜在生成モデルは通常、2段階のパイプラインを採用する。すなわち、再構成のために変分オートエンコーダーを訓練し、その後、凍結した潜在空間上で生成モデルを訓練する。再構成に最適化された潜在表現は必ずしも生成に適しているとは限らないため、両モデルを同時に訓練することは魅力的な代替手段である。しかし、直接的なエンドツーエンド訓練は、潜在空間の崩壊を引き起こしやすく、生成と再構成の間の相反する課題に直面するため、依然として困難である。我々は、異なる目的関数が潜在空間をどのように形成するかを分析することによってこの問題を再検討し、2つの重要な知見を得た。第一に、カルバック・ライブラー情報量の発散項におけるエントロピー項は、崩壊を防ぐために不可欠である。再構成と事前分布の適合は事後分布を縮小する傾向がある一方、エントロピーは非退化な潜在不確実性を維持する。第二に、再構成と生成は非対称な学習ダイナミクスを示す。再構成は高速かつ強く教師ありであるのに対し、生成はより遅く最適化が困難である。これらの知見に基づき、我々は潜在空間の崩壊を伴わない直接的なエンドツーエンド訓練を初めて達成し、生成を再構成よりも先に行う単純な戦略であるGenFirstを提案する。生成的目標は、弱い再構成圧力の下で最初に潜在空間を形成し、その後、視覚的詳細を復元するために再構成を段階的に強化する。我々はGenFirstを、厳密な尤度を持つ連続自己回帰事前分布と、暗黙的尤度を持つSiT事前分布を用いて検証する。我々のエンドツーエンド目的関数とGenFirstにより、SiTはImageNet-256においてCFGありでgFID 0.97、CFGなしで1.45を達成し、MMDiTはテキストから画像への生成においてGenEvalスコア0.90を達成する。画像生成を超えて、我々はこの枠組みを、生成と表現学習のための共有視覚潜在表現、および連続的な統合テキスト画像生成に拡張する。これらの結果は、生成事前分布とモダリティを横断した安定したエンドツーエンド潜在学習の一般性を示している。
English
Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.