ChatPaper.aiChatPaper

에너지-유도 플로우 매칭

Energy-Guided Flow Matching

August 7, 2026
저자: Haoyang Tong, Yu He, Fang Li, Lichen Ma, Jingling Fu, Dong Chen, Zhen Chen, Junshi Huang, Jie Cao
cs.AI

초록

픽셀 공간 생성 모델은 손실 있는 잠재 압축을 우회하지만, 고차원 공간에서 전역 구조와 미세한 세부 정보를 공동 학습해야 한다. 표준 플로우 매칭은 노이즈를 고정된 클린 이미지 종점으로 보간하며, 스펙트럼의 진화는 암묵적으로 학습되도록 남겨 둔다. 본 논문에서는 종점을 이동시켜 coarse-to-fine 생성 궤적을 명시적으로 모델링하는 에너지 기반 플로우 매칭(EG-FM)을 제안한다. 구체적으로, EG-FM은 고정 종점을 저주파 이미지에서 클린 이미지로 부드럽게 진화하는 열 커널 필터링 종점으로 대체한다. 이동 종점에서 고주파 신호의 비율은 이미지별 에너지 기반 스케줄링을 통해 방출되며, 이는 플로우 매칭의 속도 재설정으로 이어진다. 우리 프레임워크는 백본과 훈련 데이터의 적응이 필요 없어 훈련 및 추론 단계에서 무시할 수 있는 비용만 발생시킨다. 실험에서 EG-FM은 256×256 해상도의 ImageNet 클래스 조건부 이미지 생성 작업에서 더 적은 에폭으로 일관되게 더 낮은 FID를 달성했으며, 200 에폭에서 FID 1.55, 600 에폭에서 1.45를 기록했다. 또한 512×512 해상도 설정에서 생성 작업을 계속 훈련하여 단 40번의 고해상도 적응 에폭 만에 FID 1.58을 얻었다. 나아가 EG-FM을 텍스트-이미지 생성에 전이하여 GenEval 점수 0.85, DPG-Bench 점수 83.9를 달성했다. 코드는 https://github.com/ysng123/EG-FM에서 확인할 수 있다.
English
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256 times 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 times 512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.