用于一致性多参考图像编辑的评估-验证奖励
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
July 31, 2026
作者: Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin
cs.AI
摘要
尽管近年来的图像编辑模型取得了快速进展,多参考编辑仍然具有挑战性,尤其是在跨参考保持视觉一致性以及确保整体视觉和谐方面。强化学习已被证明对文本到图像生成和单图像编辑非常有效,但其向多参考编辑的扩展受到缺乏能捕捉多图像关系约束的合适奖励模型的阻碍。此外,直接将多模态大语言模型(MLLMs)用作零样本评估器,面临易产生幻觉的长文推理与推理能力有限的简短判断之间的关键矛盾。针对这些问题,我们提出多维评估-验证奖励(EVR)。EVR将评估分解为不同的视觉标准;对于每个标准,MLLM评估器生成多个候选假设,验证器以具体视觉证据为据核实每项声明并决定接受或拒绝,从而产生可靠且细粒度的奖励信号。结合可扩展的数据流水线,我们的方法无需改动架构即可对现成编辑器进行强化学习微调。大量实验表明,与基础Qwen-Image-Edit相比,我们的方法取得了显著提升,一致性和和谐性达到或超越了NanoBanana。
English
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.