ChatPaper.aiChatPaper

少步生成建模的感知流匹配

Perceptual Flow Matching for Few-Step Generative Modeling

July 3, 2026
作者: Chuyang Zhao, Yifei Song, Hongfa Wang, Jianlong Yuan, Yuan Zhang, Siming Fu, Zhineng Chen, Huilin Deng, Haoyang Huang, Nan Duan
cs.AI

摘要

我們提出了感知流匹配(PFM),這是一種簡單而有效的框架,用於流匹配模型中的少步生成。不同於傳統在VAE潛在空間中進行速度回歸,PFM利用預訓練的感知模型,在感知特徵空間中監督流匹配。這一簡單改變顯著提升了流匹配模型的少步生成能力,將採樣步數從35-50步減少到4-8步,同時保持生成品質。與現有的加速和蒸餾方法不同,PFM既不需要教師模型,也不需要輔助評分網絡,且能以最小改動整合至標準的流匹配訓練流程。在影像生成、視訊生成和影像編輯任務上的大量實驗證明,PFM能持續產出高品質結果,同時比現有基於蒸餾的方法產生更少的偽影。我們進一步證明,感知監督將回歸最小化器從均值尋求轉變為模式尋求,使預測偏向於流形上的模式,這些模式在粗略的少步積分下仍然準確。我們的結果揭示,當在適當的表示空間中進行監督時,標準的流匹配訓練自然能產出高品質的少步生成器。我們希望這一見解能啟發未來關於以表示為目標的高效生成建模研究。
English
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.