MeanFlowNFT:將前向過程強化學習引入平均速度生成器
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
July 16, 2026
作者: Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang
cs.AI
摘要
MeanFlow生成器通過預測時間區間內的平均速度,實現了快速的少步採樣,使其在高效生成中具備吸引力。強化學習(RL)已成為一種將擴散模型和流模型與人類偏好及任務特定目標對齊的強大方法。特別是,DiffusionNFT提供了一種高效的順向過程RL框架,無需反向過程軌跡或似然估計。然而,將此類RL方法應用於MeanFlow仍未得到充分探索。DiffusionNFT優化瞬時速度,而MeanFlow則使用平均速度進行採樣。為彌補這一差距,我們引入了MeanFlowNFT。受橋接平均速度與瞬時速度的MeanFlow恆等式的啟發,我們構建了一個誘導的瞬時速度預測器。我們將DiffusionNFT目標應用於該預測器,從而使獎勵優化對MeanFlow而言定義明確。採樣仍基於平均速度,保留了MeanFlow快速的少步生成能力。我們進一步證明MeanFlowNFT繼承了DiffusionNFT的嚴格策略改進保證。在圖像和視頻生成上的實驗表明,MeanFlowNFT持續改進基線模型。此外,它在大多數指標上(SD3.5-M上8項中的6項)優於先前的最先進RL調優少步生成器,甚至可以在僅使用少量採樣步驟的情況下超越多步RL調優擴散模型。例如,在Wan 2.1上,4步MeanFlowNFT達到84.33的VBench分數,超越了50步LongCat-Video RL(82.57)。
English
MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics (6 of 8 on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, 4-step MeanFlowNFT reaches a VBench score of 84.33, surpassing 50-step LongCat-Video RL (82.57).