AdvFD:通过对抗性Fréchet距离损失提升视觉生成
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
August 11, 2026
作者: Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang
cs.AI
摘要
弗雷歇距离最近作为一种有效的分布级目标被引入生成器后训练,补充了传统的样本级扩散和流匹配损失。然而,直接优化弗雷歇目标可能导致弗雷歇作弊。目标指标持续改善,但视觉质量和其他特征空间中的弗雷歇对齐可能停滞或恶化。我们将这一失败归因于现有弗雷歇损失所使用的静态预训练特征空间。这些特征空间提供了真实分布与生成分布之间差异的不完整且固定的视角。为了解决这一局限性,我们提出了对抗弗雷歇距离(AdvFD),它通过经校准的对抗学习表示来补充FD-Loss中的静态表示目标。AdvFD用一个可学习表示增强原始静态弗雷歇目标,该表示对抗性地最大化真实样本与生成样本之间的弗雷歇差异,同时生成器在所得自适应特征空间中最小化相同差异。为了防止对抗表示通过特征放大平凡地增加目标,我们进一步引入了真实特征白化,它归一化其尺度和协方差几何,并稳定最小-最大优化。大量实验表明,AdvFD在JiT和pMF骨干以及不同模型规模上一致地改善了一步生成器后训练。
English
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.