소수 단계 생성 모델링을 위한 지각적 흐름 매칭
Perceptual Flow Matching for Few-Step Generative Modeling
July 3, 2026
저자: Chuyang Zhao, Yifei Song, Hongfa Wang, Jianlong Yuan, Yuan Zhang, Siming Fu, Zhineng Chen, Huilin Deng, Haoyang Huang, Nan Duan
cs.AI
초록
본 논문에서는 흐름 매칭 모델의 소수 단계 생성을 위한 간단하면서도 효과적인 프레임워크인 지각 흐름 매칭(Perceptual Flow Matching, PFM)을 제안한다. 기존의 VAE 잠재 공간에서 속도 회귀를 수행하는 대신, PFM은 사전 훈련된 지각 모델을 활용하여 지각 특징 공간에서 흐름 매칭을 감독한다. 이러한 간단한 변경만으로도 흐름 매칭 모델의 소수 단계 생성 성능이 크게 향상되어, 샘플링 단계를 35~50회에서 4~8회로 줄이면서도 생성 품질을 유지한다. 기존의 가속화 및 증류 접근법과 달리, PFM은 교사 모델이나 보조 점수 네트워크가 필요하지 않으며, 최소한의 수정만으로 표준 흐름 매칭 훈련 파이프라인에 통합될 수 있다. 이미지 생성, 비디오 생성, 이미지 편집 작업에 대한 광범위한 실험을 통해 PFM이 기존 증류 기반 방법보다 적은 인공물을 생성하면서 일관되게 고품질 결과를 산출함을 입증한다. 또한, 지각 감독이 회귀 최소화기를 평균 추구에서 모드 추구로 전환시켜 예측을 다양체 상의 모드로 편향시키며, 이는 거친 소수 단계 통합 하에서도 정확성을 유지함을 보여준다. 본 연구 결과는 표준 흐름 매칭 훈련이 적절한 표현 공간에서 감독될 때 자연스럽게 고품질의 소수 단계 생성기를 도출할 수 있음을 시사한다. 이러한 통찰이 효율적인 생성 모델링을 위한 표현 인식 목적 함수에 대한 향후 연구에 영감을 주기를 기대한다.
English
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.