MeanFlowNFT: 平均速度生成器への前方過程RLの導入
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
July 16, 2026
著者: Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang
cs.AI
要旨
MeanFlow生成器は時間間隔における平均速度を予測することで高速な少数ステップサンプリングを実現し、効率的な生成において魅力的な手法である。強化学習(RL)は拡散モデルやフローモデルを人間の嗜好やタスク固有の目的に合わせるための強力な手法として台頭している。特にDiffusionNFTは、逆過程の軌跡や尤度推定を必要としない効率的な前向き過程RLフレームワークを提供する。しかし、このようなRL手法をMeanFlowに適用することは十分に検討されていない。DiffusionNFTが瞬間速度を最適化するのに対し、MeanFlowは平均速度を用いてサンプリングを行う。この乖離を埋めるために、我々はMeanFlowNFTを導入する。平均速度と瞬間速度を橋渡しするMeanFlow恒等式に着想を得て、誘導型瞬間速度予測器を構築する。我々はこの予測器にDiffusionNFTの目的関数を適用し、MeanFlowに対する報酬最適化を明確に定義する。サンプリングは引き続き平均速度に基づき、MeanFlowの高速な少数ステップ生成を維持する。さらに、MeanFlowNFTがDiffusionNFTの厳格な方策改善保証を継承することを証明する。画像および動画生成実験において、MeanFlowNFTは一貫してベースラインを改善する。さらに、先行する最先端のRL調整済み少数ステップ生成器をほとんどの指標(SD3.5-Mでは8項目中6項目)で上回り、わずかなサンプリングステップで多段階RL調整拡散モデルを凌駕することさえある。例えばWan 2.1では、4ステップのMeanFlowNFTがVBenchスコア84.33を達成し、50ステップのLongCat-Video RL(82.57)を上回っている。
English
MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics (6 of 8 on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, 4-step MeanFlowNFT reaches a VBench score of 84.33, surpassing 50-step LongCat-Video RL (82.57).