ChatPaper.aiChatPaper

MeanFlowNFT:将前向过程强化学习引入平均速度生成器

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

July 16, 2026
作者: Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang
cs.AI

摘要

MeanFlow生成器通过预测时间区间内的平均速度实现快速少步采样,使其成为高效生成的有力工具。强化学习已成为将扩散模型和流模型与人类偏好及任务特定目标对齐的强大方法。特别地,DiffusionNFT提供了一种高效的前向过程强化学习框架,无需逆过程轨迹或似然估计。然而,将此类强化学习方法应用于MeanFlow仍鲜有探索。DiffusionNFT优化瞬时速度,而MeanFlow使用平均速度进行采样。为弥合这一差距,我们提出MeanFlowNFT。受MeanFlow恒等式(桥接平均速度与瞬时速度)启发,我们构建了一个诱导的瞬时速度预测器。将DiffusionNFT目标应用于该预测器,使奖励优化对MeanFlow而言具有良好定义。采样仍基于平均速度,从而保留MeanFlow的快速少步生成特性。我们进一步证明MeanFlowNFT继承了DiffusionNFT严格的策略改进保证。在图像和视频生成上的实验表明,MeanFlowNFT持续改进基线方法。此外,它在大多数指标上(SD3.5-M上8项中的6项)优于先前的先进强化学习调优少步生成器,甚至能在仅使用少数采样步数的情况下超越多步强化学习调优扩散模型。例如,在Wan 2.1上,4步MeanFlowNFT达到84.33的VBench分数,超越了50步LongCat-Video RL(82.57)。
English
MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics (6 of 8 on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, 4-step MeanFlowNFT reaches a VBench score of 84.33, surpassing 50-step LongCat-Video RL (82.57).