GenFirst: 안정적 종단간 잠재 생성 모델링을 위한 재구성 이전 생성

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

August 29, 2026
저자: Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang
cs.AI

초록

잠재 생성 모델은 일반적으로 재구성을 위해 변분 오토인코더를 훈련한 다음, 고정된 잠재 공간에서 생성 모델을 훈련하는 2단계 파이프라인을 따른다. 재구성에 최적화된 잠재 표현이 반드시 생성에 적합한 것은 아니므로, 두 모델을 공동으로 훈련하는 것은 매력적인 대안이 된다. 그러나 직접적인 종단간 훈련은 잠재 붕괴가 발생하기 쉽고 생성-재구성 충돌에 직면하기 때문에 여전히 어려운 과제이다. 우리는 다양한 목적 함수가 잠재 공간을 어떻게 형성하는지 분석함으로써 이 문제를 재검토하고, 두 가지 핵심 통찰을 도출한다. 첫째, 쿨백-라이블러 발산 목적 함수의 엔트로피 항은 붕괴를 방지하는 데 필수적이다. 재구성과 사전 분포 적합은 사후 분포를 축소하는 경향이 있는 반면, 엔트로피는 비퇴화된 잠재 불확실성을 유지한다. 둘째, 재구성과 생성은 비대칭적 학습 동역학을 보인다. 재구성은 빠르게 진행되며 강한 지도 학습이 적용되는 반면, 생성은 더 느리고 최적화하기 어렵다. 이러한 통찰을 바탕으로 우리는 잠재 붕괴 없이 최초로 직접 종단간 훈련을 달성하고, 재구성보다 생성을 먼저 학습하는 간단한 전략인 GenFirst를 제안한다. 생성 목적 함수는 약한 재구성 압력 하에서 먼저 잠재 공간을 형성하고, 그 후 재구성이 점진적으로 강화되어 시각적 세부 정보를 복원한다. 우리는 정확한 우도를 갖는 연속 자기회귀 사전 분포와 암시적 우도를 갖는 SiT 사전 분포로 GenFirst를 검증한다. 우리의 종단간 목적 함수와 GenFirst를 사용하여 SiT는 ImageNet-256에서 CFG 사용 시 gFID 0.97, CFG 미사용 시 1.45를 달성하고, MMDiT는 텍스트-이미지 생성에서 GenEval 점수 0.90을 달성한다. 이미지 생성 외에도 우리는 이 프레임워크를 생성 및 표현 학습을 위한 공유 시각 잠재 표현과 연속 통합 텍스트-이미지 생성으로 확장한다. 이러한 결과는 다양한 생성 사전 분포와 모달리티에 걸쳐 안정적인 종단간 잠재 학습의 일반성을 입증한다.
English
Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.
PDF591September 2, 2026