ChatPaper.aiChatPaper

一致性多參考影像編輯的評估-驗證獎勵

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

July 31, 2026
作者: Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin
cs.AI

摘要

儘管近期的圖像編輯模型已取得快速進展,多參考圖像編輯仍然充滿挑戰,特別是在不同參考圖像之間維持視覺一致性,以及確保整體視覺和諧方面。強化學習已被證明在文字生成圖像與單圖像編輯上非常有效,但將其擴展至多參考編輯卻受到缺乏合適獎勵模型所阻礙,因為這類模型需能捕捉多圖像之間的關係約束。此外,單純地將多模態大型語言模型(MLLMs)用作零樣本評估器,會面臨一個關鍵矛盾:容易產生幻覺的長篇推理,與短篇判斷有限的演繹能力之間的衝突。我們提出一種多維度評估—驗證獎勵(EVR)來解決這些問題。EVR 將評估分解為不同的視覺標準;針對每一項標準,MLLM 評估器會生成多個候選假設,而驗證器則將每個主張錨定於具體視覺證據,並據以接受或否決該主張,進而產生可靠且細粒度的獎勵信號。結合可擴展的資料管線,我們的方法使得無需修改架構即可對現成編輯器進行強化學習微調。大量實驗顯示,相較於基礎的 Qwen-Image-Edit,我們的方法有顯著提升,在一致性與和諧性上達到甚至超越 NanoBanana。
English
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.