AdvFD:透過對抗弗雷歇距離損失提升視覺生成
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
August 11, 2026
作者: Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang
cs.AI
摘要
弗雷歇距離近來已成為生成器後訓練中有效的分布層級目標,補充了傳統的樣本層級擴散與流匹配損失。然而,直接最佳化弗雷歇目標可能導致弗雷歇作弊。目標指標持續改善,但其他特徵空間中的視覺品質與弗雷歇對齊可能停滯甚至惡化。我們將此失敗歸因於現有弗雷歇損失所使用的靜態預訓練特徵空間。這些特徵空間提供了不完整且固定的真實分布與生成分布之間差異的視角。為了解決此限制,我們提出了對抗式弗雷歇距離(AdvFD),以經校準的對抗式學習表徵補充 FD-Loss 中的靜態表徵目標。AdvFD 以可學習表徵增強原始靜態弗雷歇目標,該表徵對抗性地最大化真實樣本與生成樣本之間的弗雷歇差異,同時生成器在所得的自適應特徵空間中最小化相同的差異。為防止對抗式表徵透過特徵放大平凡地增加目標,我們進一步引入真實特徵白化,以正規化其尺度與共變異數幾何,並穩定最小—最大最佳化。大量實驗顯示,AdvFD 在 JiT 與 pMF 兩種骨幹網路以及不同模型規模上,皆一致地改善單步生成器後訓練。
English
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.