少ステップ生成モデリングのための知覚的フローマッチング
Perceptual Flow Matching for Few-Step Generative Modeling
July 3, 2026
著者: Chuyang Zhao, Yifei Song, Hongfa Wang, Jianlong Yuan, Yuan Zhang, Siming Fu, Zhineng Chen, Huilin Deng, Haoyang Huang, Nan Duan
cs.AI
要旨
我々は、フローマッチングモデルにおける数ステップ生成のためのシンプルかつ効果的なフレームワークである、知覚的フローマッチング(PFM)を提案する。従来のVAE潜在空間での速度回帰ではなく、PFMは事前学習された知覚モデルを用いて知覚的特徴空間におけるフローマッチングを教師する。この単純な変更により、フローマッチングモデルの数ステップ生成能力が大幅に向上し、サンプリングステップ数を35〜50から4〜8に削減しながら生成品質を維持する。既存の高速化や蒸留の手法とは異なり、PFMは教師モデルや補助スコアネットワークを必要とせず、最小限の修正で標準的なフローマッチング学習パイプラインに統合できる。画像生成、動画生成、画像編集タスクにおける広範な実験により、PFMは既存の蒸留ベースの手法よりもアーティファクトが少なく、一貫して高品質な結果を生み出すことを示す。さらに、知覚的教師が回帰の最小化を平均探索からモード探索へと移行させ、粗い数ステップ統合でも正確さを保つオン・マニフォールドモードへの予測の偏りを生むことを明らかにする。この結果は、標準的なフローマッチング学習が適切な表現空間で教師されることで、自然に高品質な数ステップ生成器を生み出せることを示している。この洞察が、効率的な生成モデリングのための表現認識型目的関数に関する今後の研究を促進することを期待する。
English
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.